SRE - DataPlatform
ParisVeepeeTech – Site Reliability Engineering /Permanent /HybridJoin a transversal SRE community embedded in a product-oriented Data Platform team of 40–50 engineers, analysts, and data scientists across France and Spain. You'll drive the reliability and scalability of a next-generation Lakehouse platform — anchored on Trino, Iceberg, and on-prem object storage — while leading the transition from public cloud to a resilient hybrid/on-prem architecture.ð TASKS Platform Reliability & SRE foundations Own reliability of core data services: Trino, Iceberg, S3 / Ceph, Kafka, Kafka Connect, Schema RegistryDefine and enforce SLIs/SLOs, error budgets, and on-call runbooks — solid SRE foundations are non-negotiableBuild full-stack observability with Prometheus and Grafana: metrics, dashboards, alerting pipelines, and anomaly detectionManage and harden PostgreSQL clusters via Patroni for high-availability control-plane servicesKafka ecosystem — Connect & Schema governanceOperate and scale Kafka Connect clusters: connector lifecycle, offset management, dead-letter queues, and task rebalancingMaintain the Schema Registry as the single source of truth for Avro/Protobuf/JSON schemas — enforce compatibility rules and schema evolution policiesMonitor consumer lag, connector throughput, and broker health via Prometheus JMX exporters and Grafana dashboardsEnsure end-to-end data contract integrity between producers and Iceberg/S3 consumersKubernetes, Kube-in-Kube & CrossplaneOperate production Kubernetes clusters (GKE/EKS + on-prem) — capacity planning, upgrades, PodDisruptionBudgets, resource quotasArchitect and manage Kube-in-Kube topologies to provide strong tenant isolation for data platform workloads — each team gets a dedicated virtual cluster without the overhead of a full physical clusterAutomate infrastructure and resource provisioning with Crossplane: define composite resources (XRDs) so data teams can self-serve Kafka topics, Trino namespaces, and S3 buckets through Kubernetes-native APIsMaintain GitOps pipelines for platform deployment and configuration drift detectionLakehouse architecture & cloud migrationMigrate from public cloud data warehouse to VeepeeCloud Iceberg-based lakehouse — managing coexistence, schema evolution, and time-travelArchitect resilient ingestion, transformation, and serving layers around Trino + S3Optimize Trino query performance: memory limits, spilling, cost-based optimizer tuningAgentic & developer enablementBuild agentic self-service tooling so data teams can provision Trino/Iceberg resources and Kafka Connect pipelines autonomously via Crossplane — reducing toil and ops bottlenecksDevelop FinOps dashboards (compute, storage, query cost) with Grafana and Prometheus-based cost exportersWrite clear technical documentation, runbooks, and internal ADRsMulti-DC resilience & DRPDesign and implement multi-datacenter strategies across FR1 / NL1 — active-active and active-passive topologiesLeverage Fast Erasure Coding on object storage (Ceph/S3) to maximize durability with minimal replication overheadEnsure data replication consistency across sites for Iceberg table metadata, Trino catalogs, and Schema Registry subjectsLead DRP exercises: failover playbooks, RTO/RPO validation, postmortems ð MUST HAVE skills Must haveStrong experience with Kubernetes in production environmentsExperience with Kube-in-Kube technologies (vCluster or similar)Solid understanding of SRE principles (SLIs/SLOs, error budgets)Experience with Prometheus and GrafanaExperience with Infrastructure as Code (Terraform or similar)Experience with CrossplaneFamiliarity with GitOps workflowsExperience with S3 and object storage technologiesExperience with PostgreSQL and PatroniExperience with Kafka, Kafka Connect, and Schema RegistryFluent in Englishð NICE TO HAVE skillsExperience with multi-datacenter architectures (FR1/NL1)Experience designing disaster recovery plans and failover playbooksExperience with Fast Erasure Coding (Ceph/S3)Experience with Trino, Iceberg, and Lakehouse technologiesExperience with AirflowExperience building agentic self-service platformsKnowledge of FinOps and cost optimization practicesProgramming experience in Python, Java, or Go BENEFITSVariable bonusE-learning platform (self-education courses)Meetups & conferences (local and international)Flexible office — up to 2 days remoteInternational teams (France & Spain)️ RECRUITMENT PROCESS1️⃣ 30-minute HR Screen with a Veepeeᵀᵉᶜʰ Recruiter2️⃣ General Technical exchange3️⃣ Technical exchange with the manager4️⃣ Team InterviewWe are convinced that it is up to you to define the way you work, to develop yourself and to progress. At Veepee we guarantee that you can just be yourself!For the service of diversity and inclusion, Veepee is committed to reviewing all applications received on an equal basis. ðCOMPANY For more information about our ecosystem : We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Emplois Recommandés
Account Executive - CDI
Salaire: €100k OTE Account Executive (H/F) | Profil chasseur Le contexte Sparkway accompagne l'un de ses clients, un acteur établi et reconnu sur son marché, dans le renforcement de son équipe c…
Ingénieur Qualité Produit & Maintenance N3 H/F - Stage
Type de contrat : Stage de fin d’études Localisation : Paris, 7ème (présentiel) Date de début : septembre 2026 Expérience : Idéalement un stage réalisé auparavant A propos de NW NW v…
Architecte d'entreprise confirmé F/H
Détails de l'offre Famille de métiers Informatique et système d'information Architecture, dév…
Dermatologue CDI Cadre - Centre Médical & Dentaire - Paris 15e
Le Centre Médical et Dentaire Vaugirard MGEN - - recherche - Un Dermatologue (H/F) - Poste en CDI, statut cadre - Temps plein ou partiel - - - Présentation du Site ou de la Direction - Le Centre Médic…
Responsable commercial(e )- CDD 5mois
This position is at REED Global The selection process will be fully managed by REED Global. This opportunity is available in Paris - France, Puteaux - France. -- Poste en CDD 5 mois Le mond…
Facility Manager H/F
Le poste de Facility Manager H/F En tant que FACILITY MANAGER H/F, vos responsabilités consisteront à : * Superviser l'entretien et la maintenance des installations, en garantissant leur bon …
Chargé de projet data et automatisation(H/F)
Type de recrutement : Stage Domaine de compétences :Donnée Ville :Paris Département :Paris Description du poste: Au sein de la Direction des Entreprises (DE), le service OSMOSE regroupe près de…
Un gestionnaire développement des compétences (H/F)
Détails de l'offre Famille de métiers Développement économique et emploi Politiques d'emploi,…
Sr. Devops Engineer II
Who we areDoubleVerify is an Israeli-founded big data analytics company (Stock: NYSE: DV). We track and analyze tens of billions of ads every day for the biggest brands in the world.We operate at a ma…
Apprenti chargé de projet médico-social — H/F
Apprenti chargé de projet médico-social — H/F Nous recherchons un(e) apprenti(e) chargé(e) de projet médico-social pour rejoindre notre équipe au siège et collaborer avec nos établissements en Île-de…