PhD Position F/M Visualization of the Plausibility and Bias for Data Resources used in a Geographic Digital Twin
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
En continuant, vous acceptez nos Conditions d’utilisation & Politique de confidentialité.
The Inria Saclay-Île-de-France Research Centre was established in 2008. It has developed as part of the Saclay site in partnership with Paris-Saclay University and with the Institut Polytechnique de Paris .
The centre has 39 project teams, 27 of which operate jointly with Paris-Saclay University and the Institut Polytechnique de Paris; Its activities occupy over 600 people, scientists and research and innovation support staff, including 44 different nationalities.
The project will be hosted by Inria’s AVIZ team and supervised by : Maria Lobo-Gunther
This advertised topic targets students focused on visualization and data analysis, with a background in artificial intelligence. It is not focused purely on AI. Also, please note that we only consider applications sent in English.
In the field of visualization, questions on the visualization of data uncertainty have been a core part of the research to date [e.g., 1]. Yet past work largely often silently considered data uncertainty to arise (primarily) from measurement uncertainty or as connected to some statistical analysis of captured data values. Only relatively recently have questions of implicit notions of error [4] or data hunches [3] entered the discussions of this general problem—these cover aspects of data imprecision that can arise from various sources: e.g., different forms of measurement or data recording, biases in how people access data, and even intentionally introduced errors. For example, spatial geographic data may not only come from systematically controlled sources but also from, for instance, from multiple sets of data that come from different institutions with different data collection policies, from contributions from the general public, or even by sourcing data from social media. Naturally, this introduces a variety of different levels of data quality and plausibility for spatial data. Sometimes experts are aware of these data caveats (that exist even for professionally collected datasets) and we can try to visualize them [3–5], yet in other cases the data collection happens in a way that this knowledge is not available—e.g., when data is retro-actively collected from sources that were not (primarily) created for recording this data in the first place—and we have to find ways of recovering such information retroactively and without access to the ground truth [2]. For instance, image collection sites such as Flickr partially show geo-located images, from which locations of certain points of interest can be extracted. In this specific example we ourselves recovered species distribution data from images posted on photo sharing platforms, and found that such geographic data is subject to many biases and errors. Yet this data analysis so far [2] has been a manual process and also only recovers anecdotal information about the existence of data errors and biases. In this PhD research project we will investigate if those results generalize to other kinds of spatial data (e.g., automatically detected features in remote sensing imagery, other volunteered geographic information, other forms of 2D and 3D spatial data). Then we will investigate ways to automate the error and bias identification process as well as to quantify the existence of such biases and errors in some form of plausibility measure, to be visualized in the digital twin [6].
The funding for this thesis comes from the EDT project (https://edtlab.fr/en/).
Further informations : https://www.edtlab.fr/en/join-us/phd-pc5-phd2-estimatingbiases
- analysis and visualization of the plausibility of crowd-sourced and/or heterogenous geographic data
- * use of artificial intelligence to automate the analysis of the data errors and biases
- Digital twins [6] are virtual representations of real-world products, systems, or processes, enabling simulation, integration, testing, monitoring, and maintenance. They play a pivotal role in optimizing complex systems across a wide range of domains, from industrial manufacturing and energy to environmental monitoring and healthcare.
- The intellectual property is shared among all collaborators and the scientific results will be published in international journals and/or at
- international conferences.
- For the remaining context please see the abstract.
Method :
Typical requirements analysis-design-implementation-evaluation approach that is used in most of HCI and visualization.
Expected result :
Prototypical implementations in the context and within the framework of the EDT project (https://edtlab.fr/en/), study results, scientific publications.
&l