Français Anglais
Accueil Annuaire Plan du site
Accueil > Evenements > Séminaires
Séminaire d'équipe(s) LaHDAK
Beyond declarative mapping and cleaning
Paolo Papotti

02 February 2015, 14h00 - 02 February 2015, 16h00
Salle/Bat : 445/PCRI-N
Contact : ppapotti@qf.org.qa

Activités de recherche : Algorithmes pour les grands volumes de données distribuées

Résumé :
In the "big data" era, data integration is a popular activity both in
academia and in industry. Integrating hundreds of heterogeneous
sources on a daily basis requires a great amount of manual work in
order to have data that is polished enough to be useful in the final
applications, such as querying and mining. The problem is ever harder in practice, as data is often dirty in nature because of typos, duplicates, and so on, that can lead to poor results in the analytic tasks.

Over the last ten years, several successful systems have been proposed to tackle this challenge with a formal, declarative approach based on first order logic. However, despite the positive results, there is still a gap between these proposals and the leading commercial systems. The latter are harder to maintain, to debug, and to test, but provide the level of personalization and detail that are needed to solve “real-world” problems. In this talk, I will describe some of my results in tackling mapping and cleaning with a declarative approach, and how this experience has pushed me to explore a new way that can take the best of both worlds.

Short Bio:

Paolo Papotti is a scientist in the Data Analytics center at Qatar
Computing Research Institute (QCRI). He holds a Ph.D degree in
computer science from Roma Tre University (Italy, 2007), where he also was Assistant Professor before joining QCRI. He had visiting appointments at IBM Almaden (USA) and at the UC Santa Cruz (USA). His research topics are in the general area of information integration and data quality.

Pour en savoir plus : http://www.qcri.qa/our-people/bio?pid=34&name=PaoloPapotti
Séminaires
Heterogeneous Treatment Effects Estimation: When M
Raisonnement automatique
Thursday 02 June 2022 - 10h30
Salle : 2011 - DIG-Moulon
Naoufal Acharki .............................................

Witness Generation for JSON Schema
Langages et systèmes centrés données
Monday 30 May 2022 - 00h00
Salle : 455 - PCRI-N
Mohamed-Amine BAAZIZI .............................................

TUTORIAL CODALAB - Apprenez à organiser un challen
Wednesday 13 April 2022 - 00h00
Salle : 1 - DIG-Moulon
Adrien Pavao .............................................

Generative Neural Networks for Observational Causa
Raisonnement automatique
Thursday 07 April 2022 - 10h30
Salle : 2011 - DIG-Moulon
Diviyan Kalainathan .............................................

Datamining in Epi- and Phylogenetics
Tuesday 15 March 2022 - 11h00
Salle : 455 - PCRI-N
Thomas Haschka .............................................