Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

13 July 2015

Le traitement automatique des langues (enfin) à l’honneur


Confrontée à un volume d’information toujours croissant, l’Europe découvre, ravie, la valeur du traitement automatique des langues, ciment de la construction européenne.

Lors du récent sommet LT-Innovate, Alexander De Croo, vice premier ministre de Belgique et ministre de l’Economie digitale ainsi que Robert Madelin, directeur général de la DG CONNECT à la Commission européenne, ont envoyé un message très clair à la communauté du Traitement Automatique des Langues (TAL) : « Nous comprenons aujourd’hui l’importance de votre discipline et le rôle qu’elle joue dans le développement économique de l’Europe. Nous apprécions aussi votre capacité à transformer et adapter votre discours à nos préoccupations économique et politique ».

Ce message, illustré dans les interventions régulières des intervenants politiques, souligne la prise de conscience du rôle fondamental du traitement automatique des langues.

Cette « compréhension déclarée » serait ainsi liée à la transformation du discours de notre discipline envers les autorités. Je n’en suis pas aussi convaincu que cela. Je n’ai pas le sentiment que notre discours ai subitement ou progressivement changé fondamentalement, et ce quelle que soit la discipline concernée, la traduction, la reconnaissance vocale ou encore l’analyse sémantique.
Cette soudaine prise de conscience des autorités européenne me semble être davantage une conséquence de leur difficulté, voire de leur impossibilité à faire face au volume d’information qui les submerge aujourd’hui.

En ce sens, nous pouvons rappeler à la communauté et aux autorités, que l’un des 3 V du Big Data – la Variété - caractérise intrinsèquement la masse des données qu’il faut appréhender et traiter.
C’est bien cet enfant de l’ère numérique qui a éveillé les consciences sur l’importance des données, leur nature, leur diversité, leur masse, pour l’aide à la décision technique, économique et politique.
Cependant, quelles qu’en soient les raisons, cette prise de conscience, dans le contexte du « Digital Agenda for Europe » est une excellente nouvelle pour notre communauté. Celle-ci, rappelons-le, est composée à la fois d’universitaires, mais également d’un grand nombre de PME à travers toute l’Europe. Il apparaît donc aujourd’hui que nous sommes clairement identifiés et reconnus pour nos expertises variées et notre valeur contributive au développement et aux enjeux européens.

Nous le savons, et j’ai pu le vérifier lors de notre réunion annuelle, toutes les entreprises engagées de notre communauté connaissent bien la manière dont elles peuvent contribuer à ce développement stratégique. En revanche, trop nombreuses sont celles qui finissent par baisser les bras au moment de se confronter aux mécanismes administratifs complexes et statutaires de l’Europe. Nous sommes en général des entreprises de petite taille et malheureusement pas toujours correctement équipées pour échanger d’égal à égal avec les autorités Européennes à l’occasion de projets de type H2020 ou autre. Nous avons parfois le sentiment regrettable qu’au cours des 15 dernières années, le fossé entre nous continue inexorablement de se creuser, qu’une communication simple et directe reste toujours difficile et qu’au final, l’Europe ne sait pas nous accompagner.

Il est urgent et impératif que l’Europe assume et entretienne un rôle d’accompagnement stratégique – à l’instar des États-Unis – auprès de nos PME innovantes, de nos start-up, parfois fragiles, afin d’assurer des perspectives de développement pérenne à moyen et long terme. L’Europe doit comprendre l’importance stratégique des technologies innovantes que nous développons pour servir, entre autre, l’indépendance technologique, économique, culturelle et juridique de notre continent européen.

Il est heureux que nos représentants européens prennent conscience de notre existence technologique et de notre valeur associée. Il est temps maintenant que notre Europe administrative se mette à notre hauteur afin de nous apporter une aide active en nous impliquant dans des projets d’exécution et de production. L’un des premiers bénéfices attendus permettrait certainement de simplifier et d’optimiser ses propres rouages administratifs...

Charles Huot est le directeur général délégué et co-fondateur de TEMIS, une société de gestion des données non structurées. TEMIS aide les entreprises à archiver, gérer, analyser, trouver et partager un volume d’informations toujours croissant. Cet article a été publié aussi sur EurActiv.

10 June 2015

Major disruption ahead in the language industry!


http://www.gala-global.org/GALAxy/Q2-2015/5665
The Q2 issue of GALAxy, the quarterly newsletter of our partner association GALA, is guest edited by LT-Innovate Chairman Jochen Hummel (@JochenHummel) with a thought provoking piece on how Language Technology will leverage Big Data and transform the industry.

Several other articles are contributed and/or co-authored by LT-Innovate members and partners:
  • Big Data and the Translation Industry: Three Technology Challenges by Andrew Joscelyne, LT Innovate
  • Finding New Business Segments Through Big Data by Michael Wetzel, Coreon GmbH & Matthias Heyn, SDL plc
  • How to Improve Your Relationship with Machine Translation co-authored by Heidi Depraetere, CrossLang
  • Unlocking Language Resource Assets by Christian Galinski, Infoterm
  • Riga Summit Forges a Unified Vision for Multilingual Europe by Rihards Kalniņš, Tilde

22 May 2014

Language Technology is the drill to make Big Data "oil" flow in Europe!



It has always surprised me how much is written and talked about Big Data without pointing to the main barrier to the data revolution: our many languages (more than 60 in Europe alone). The numbers surely differ from sector to sector, but a fair guess would be that half of big data is unstructured, i.e. text. Most multimedia data is also converted to text (speech-to-text, tagging, metadata) before further processing. Text in Europe is always multilingual.

Europe prides itself of an “undeniable competitive advantage, thanks to [its] computer literacy level”. In fact, we have had this advantage for decades, but so far it hasn’t helped much. Good brains and companies are systematically bought by our American friends. No, we rather have to focus on what is specific for Europe. On what we have and the US doesn’t. Maybe even if it is a disadvantage - at first sight.

What makes Europe special and different is the fact that we are trying to build a Single Market in spite of our different cultures and systems. Our multilingualism is always seen as a challenge, a big disadvantage. Most Big Data applications only work well in English and, with some luck, okayish in German, Spanish, or French. Smaller EU countries with lesser spoken languages are basically excluded from the data revolution. The dominance of English in content and tools is the reason for the US lead in Big Data. Many European companies have reacted to this and now use English as their corporate language. But Big Data is often big because it originates from customers and citizens. And these rather use their own languages.

What if we managed to turn this perceived handicap of a multilingual Europe into an asset? Overcoming the language barriers would be a great step towards a Single Market. We would make sure that smaller Member States participate and perhaps become drivers of the data revolution. Even more importantly, Europe would become the fittest for the global markets. The BRICs and all other emerging economies do not accept any more the dominance of English. Europe has a unique chance... if it solves a problem the Americans do not have, or discover too late.

The real opportunity is therefore to create the Digital Single Market for content/data independently of the latter's (linguistic) origin. This would require that we overcome the language-silos in which most data remains captive and make all data language-neutral.

To achieve this, we urgently need a European Language Cloud. For all text based Big Data applications the European Language Cloud is a web-based set of APIs that provides the basic functionality to build products for all languages of the Single Digital Markets and Europe’s main trading partners. For more information, see my previous post.

While the European language technology industry might not have all the solutions readily available to deliver the European Language Cloud, many language resources could be pooled as a first step. In addition, many technologies are presently entering into a phase of maturity (after decades of European investment into R&D) and could be harnessed - through a set of common APIs - into a viral Language Infrastructure. This would go a long way towards delivering the European Language Cloud... without which the Big Data oil will only continue to flow from English grounds.

Jochen Hummel
CEO, ESTeam AB - Chairman, LT-Innovate

07 October 2013

Actuate and Bitext Announce Collaboration to Deliver Text Analytics Engines and Sentiment Analysis for Big Data through BIRT

Actuate Corporation, The BIRT Company delivering more insights to more people than all BI companies combined, today announced their cooperation with Bitext, in parallel with Bitext’s U.S. event in San Francisco this evening at WeWork. Bitext provides text analytics engines – inherently multilingual semantic technologies including text analytics and natural language interfaces – sporting one of the highest degrees of accuracy available today. Bitext recently announced a partnership with Salesforce.com as well.

Combined with Actuate’s BIRT iHub™ development tools and platform, or with Actuate’s BIRT Analytics™ 4.2 predictive analysis solution, Bitext provides two main advantages for the BIRT developer or end user: it produces highly accurate precision and recall; and it lends itself easily to a development process based on continuous improvement. BIRT Analytics 4.2 and Bitext are now available as a combined solution from Actuate.

We are very pleased to be working together with Actuate to further enrich their leading BIRT commercial suite with Bitext text and semantic analysis power,” said Antonio Valderrabanos, CEO and Founder, Bitext. “With Bitext analyzing unstructured data words as well as meaning, and Actuate performing advanced analysis of structured as well as unstructured, we cover the world of data.
Bitext enables entity and concept extraction, categorization, and sentiment analysis with a focus on customer-centric business areas such as marketing; customer relationship management and support; content analytics; and any line of business unit that requires advanced analytics. Examples of solutions include text analytics (entity extraction, concept extraction, and sentiment analysis), metatagging (enhanced indexing) and search (natural language interfaces). Currently available for 10 languages, Bitext enables the addition of new languages by including new data sources (dictionaries and grammatical rules).

Our collaboration with Bitext – providers of advanced semantic solutions for social media, search, and more – extends the types of analysis that can be performed with Actuate’s commercial BIRT developer and end-user platform or solution, by adding the ability to score sentiment toward products and services,” said Josep Arroyo, VP of Analytic Solutions at Actuate. “Users of Actuate with Bitext can now tap more than just negative or positive sentiment analysis. They can also visualize anticipated risks, opportunities and threats for personalized insights, in a single display on any device.

For a demo of BIRT Analytics 4.2, please visit Actuate’s YouTube Channel
For a demo of Bitext’s new API, please visit Bitext website

30 January 2013

European Data Forum: Free Registration and Call for Contribution are open!

The European Data Forum (EDF) 2013 takes place on April 9-10, 2013 in Dublin (Ireland). It is the annual meeting-point for data practitioners from industry, research, the public-sector and the community, to discuss the opportunities and challenges of the emerging Big Data Economy in Europe.
EDF aims to bring together all stakeholders involved in the data value chain to exchange ideas and develop actionable roadmaps addressing these challenges and opportunities in order to strengthen the European data economy and its positioning worldwide.


Event if this event is without sign-up fees, we ask for registration for planning purposes

The deadline for submissions is: 22nd Feb 2013, 02.00pm CET

05 September 2012

Semantic Technology’s Role in Big Data Solutions?

Forbes has published an article that points out an opportunity for Semantic Technology companies. The article discusses the lack of understanding in companies around big data. The author, , concludes:

"Ready or not, the IT department can expect the business units to come back begging for help, just as they did in the 1990s. And even before business users start crawling back, I can imagine the red faces of all those IT executives that worked so hard in the last five years or so on establishing a “data governance” model for their companies and putting together master data management policies. The Garner analysts warned their listeners that big data “could break existing governance models.” This is similar to what Forrester analyst Boris Evelson wrote earlier this month: “You may find that all of your best DW, BI, MDM practices for SDLC, PMO and Governance aren’t directly applicable to or just don’t work for Big Data. This is where the real challenge of Big Data currently lies. I personally have not seen a good example of best practices around managing and governing Big Data. If you have one, I’d love to see it!”"