DATA COLLECTION
Il y a 4 semaines
Paris, Ile-de-France
STATION F
Temps plein
Gratuit avec email ou Google
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Gratuit avec email ou Google
En continuant, vous acceptez nos Conditions d’utilisation & Politique de confidentialité.
About
DATA&DATA helps luxury brands understand online market dynamics. We aggregate and analyze large-scale data from across the web to provide actionable insights into the pricing, availability, and visibility of high-end consumer goods.
Our core asset is our data: collected at scale from hundreds of sources, checked, normalized, and turned into analytics for some of the world's most iconic luxury brands. We're a small team, which means data collection work here is concrete and immediately useful.
Job Description
As a Web Scraping & Data Collection Intern, you'll be responsible for feeding and monitoring the data that powers our platform. Concretely, you'll:
• Configure and run scrapers on new and existing websites using our in-house scraping framework
• Investigate target websites to identify undocumented or public APIs, and figure out the most reliable way to collect their data
• Write Python scripts and notebooks to collect, extract, and structure data from a wide variety of sources
• Run quality checks on collected data: completeness, consistency, detecting when a source breaks or silently changes
• Run and monitor existing collection pipelines, and flag or fix issues when a source stops behaving
• Answer ad hoc questions on our database with SQL queries (volumes, coverage, anomalies) for internal and client-facing needs Preferred Experience You must be enrolled in a school or university able to provide a convention de stage for the full 6 months. Must-haves
• Working knowledge of Python: requests, playwright (or Selenium), plus the basics (loops, functions, files, JSON, virtual environments)
• Basic SQL: you can write a SELECT with a WHERE, a GROUP BY, and a simple JOIN without needing to look everything up
• Comfort reading HTML and using browser dev tools (inspecting the DOM, reading the network tab)
• Git basics for version control
• Autonomy: you're comfortable investigating a problem on your own before asking
• Fluency in English (written & spoken); French is a bonus Nice-to-haves
• Prior experience with web scraping, in any context (personal projects count)
• Familiarity with anti-bot mechanisms, proxies, or headless browsers
• Experience with pandas for quick data checks What You Get
• Real, messy, large-scale data from day one — the kind you can't get from a course project
• Genuine technical depth on scraping: reverse-engineering APIs, dealing with sites that don't want to be scraped, keeping collection reliable at scale
• Autonomy and ownership over the sources you handle
• A flat structure, flexible hours, casual dress, no bureaucracy Recruitment Process
• Initial screening – We review your resume and any additional materials you submit
• Phone interview – A call to discuss your background, your motivation, and a few basic technical questions That's it. No take-home assignment, no multi-round process. Additional Information
• Contract Type: Internship (Between 5 and 7 months)
•
Location:
Paris
• Education Level: Bachelor's Degree
• Occasional remote authorized
Job Description
As a Web Scraping & Data Collection Intern, you'll be responsible for feeding and monitoring the data that powers our platform. Concretely, you'll:
• Configure and run scrapers on new and existing websites using our in-house scraping framework
• Investigate target websites to identify undocumented or public APIs, and figure out the most reliable way to collect their data
• Write Python scripts and notebooks to collect, extract, and structure data from a wide variety of sources
• Run quality checks on collected data: completeness, consistency, detecting when a source breaks or silently changes
• Run and monitor existing collection pipelines, and flag or fix issues when a source stops behaving
• Answer ad hoc questions on our database with SQL queries (volumes, coverage, anomalies) for internal and client-facing needs Preferred Experience You must be enrolled in a school or university able to provide a convention de stage for the full 6 months. Must-haves
• Working knowledge of Python: requests, playwright (or Selenium), plus the basics (loops, functions, files, JSON, virtual environments)
• Basic SQL: you can write a SELECT with a WHERE, a GROUP BY, and a simple JOIN without needing to look everything up
• Comfort reading HTML and using browser dev tools (inspecting the DOM, reading the network tab)
• Git basics for version control
• Autonomy: you're comfortable investigating a problem on your own before asking
• Fluency in English (written & spoken); French is a bonus Nice-to-haves
• Prior experience with web scraping, in any context (personal projects count)
• Familiarity with anti-bot mechanisms, proxies, or headless browsers
• Experience with pandas for quick data checks What You Get
• Real, messy, large-scale data from day one — the kind you can't get from a course project
• Genuine technical depth on scraping: reverse-engineering APIs, dealing with sites that don't want to be scraped, keeping collection reliable at scale
• Autonomy and ownership over the sources you handle
• A flat structure, flexible hours, casual dress, no bureaucracy Recruitment Process
• Initial screening – We review your resume and any additional materials you submit
• Phone interview – A call to discuss your background, your motivation, and a few basic technical questions That's it. No take-home assignment, no multi-round process. Additional Information
• Contract Type: Internship (Between 5 and 7 months)
•
Location:
Paris
• Education Level: Bachelor's Degree
• Occasional remote authorized