Software Engineer

Datadog

Full-Time   |   US - Onsite (Boston, Lisbon, Madrid, NYC, Paris, Tel Aviv)

Company Overview

Datadog is a global leader in cloud‑native observability, delivering a unified SaaS platform that consolidates infrastructure monitoring, application performance monitoring (APM), log management, synthetic monitoring, security analytics, and AI‑driven telemetry for enterprises of every size. Since its founding in 2010, Datadog has grown to 6,500+ employees and serves tens of thousands of customers across technology, finance, healthcare, media, and e‑commerce sectors. The platform ingests, stores, and serves millions of events per second—metrics, traces, and logs—across a multi‑region, multi‑cloud architecture that spans AWS, Azure, and Google Cloud. Its core data layer, known as Husky, is a columnar, distributed storage system that supports low‑latency analytics, exactly‑once semantics, and automatic sharding to meet the demands of large, dynamic tenants. The engineering culture prioritizes rapid delivery, automated testing, rigorous code reviews, and continuous improvement, enabling teams to ship features at the speed demanded by modern DevOps pipelines.

Datadog’s product portfolio includes a breadth of monitoring signals: from traditional system metrics and container health checks to network performance, browser experience, serverless functions, and synthetic testing. Customers leverage the platform for real‑time incident detection, root‑cause analysis, capacity planning, and security posture management. The data ecosystem extends into the observability stack, providing a single API to query all telemetry, embed dashboards, and trigger alerts. This holistic view is supported by a robust infrastructure that includes Kafka queues, Kubernetes clusters, Redis caches, and a highly scalable metadata store built on FoundationDB, ensuring strong consistency across the platform.

Culture at Datadog is defined by transparency, inclusivity, and a bias toward action. The organization operates with a hybrid model that offers onsite offices in Boston, Lisbon, Madrid, New York City, Paris, and Tel Aviv, alongside a strong remote‑first mentality. Employees benefit from unlimited paid time off, comprehensive health benefits, equity packages, learning stipends, and a work‑from‑anywhere stipend that supports a healthy work‑life balance. Diversity, equity, and inclusion initiatives are embedded in hiring, promotion, and partner programs, fostering an environment where every voice can contribute to shaping the next generation of observability.

Key Responsibilities

  • Architect Ingestion Pipelines: Design and extend distributed ingestion micro‑services capable of processing 10⁶+ events per second, ensuring exactly‑once semantics across multi‑region Kafka clusters. Develop data models and sharding strategies that balance load, reduce skew, and minimize cold starts for new tenants.
  • Enhance the Husky Storage Engine: Own the next‑generation columnar storage layer—including schema evolution, compaction algorithms, and metadata consistency. Work with data engineers to incorporate new event types such as network logs, user‑experience traces, and security telemetry, ensuring they can be queried with sub‑second latency.
  • Optimize Observability Telemetry: Collaborate with the Tracing team to reduce ingest latency, improve span correlation and context propagation, and support distributed tracing across micro‑service architectures. Implement sampling strategies that balance cardinality and accuracy for high‑volume applications.
  • Feature Design & Delivery: Lead the design and implementation of platform APIs that expose real‑time dashboards, alerting rules, and synthetic monitoring controls. Translate product roadmaps into scalable services that can be safely rolled out across millions of tenants.
  • Reliability & Security: Define and enforce sharding placement, placement rotation, and backpressure thresholds to prevent cascading failures during traffic spikes. Implement role‑based access controls, audit logging, and encryption at rest/in transit across the data layer.
  • Mentorship & Culture: Coach junior engineers, conduct technical reviews, and promote a culture of continuous learning. Champion best practices—clean architecture, test‑driven development, CI/CD pipelines, and automated observability—across cross‑functional teams.

Required Skills

  • Distributed Systems & Concurrency: At least 8+ years of experience designing high‑throughput, fault‑tolerant services written primarily in Go, with additional proficiency in Java or Python. Deep understanding of concurrency primitives, synchronization, lock‑free algorithms, and event sourcing patterns.
  • Data Ingestion & Storage: Hands‑on experience building ingestion pipelines using Kafka, Pulsar, or similar event streaming platforms. Proven expertise in designing sharding and placement algorithms that guarantee exactly‑once processing and low‑latency querying in columnar storage or time‑series databases.
  • Cloud & Container Orchestration: Strong background deploying and scaling services on Kubernetes clusters in AWS, Azure, or Google Cloud. Ability to leverage cloud‑native services (e.g., EKS, GKE, AKS, CloudRun) and to design autoscaling, load‑balancing, and resilience strategies.
  • Observability Domain Knowledge: Familiarity with metrics, traces, logs, synthetic monitoring, and security analytics. Knowledge of distributed tracing tools (OpenTelemetry, Zipkin, Jaeger) and metrics aggregation techniques.
  • System Design & Architecture: Ability to architect large‑scale systems that span multiple services, ensuring performance, consistency, and scalability. Experience with API design, gRPC/REST, and protocol buffers.
  • Software Engineering Practices: Strong emphasis on code quality—clean architecture, unit/integration testing, static analysis, CI/CD pipelines, and automated release processes. Experience with monitoring, alerting, and blameless post‑mortems.

Ideal Candidate

  • Demonstrates a passion for building reliable, distributed systems that form the foundation for observability at scale.
  • Holds 5–10 years of engineering experience in a high‑traffic platform, with a proven track record of delivering services that handle millions of events per second, handling traffic surges and multi‑region deployments.
  • Excels in Go, but can also lead teams working in Java or Python. Comfortable writing production‑grade code, writing tests, and participating in code reviews.
  • Thrives in a collaborative, fast‑paced environment, willing to experiment with new architectures, refactor legacy components, and guide peers toward best practices.
  • Communicates complex technical concepts clearly to cross‑functional stakeholders—product managers, designers, and executive leaders.
  • Seeks continuous learning: explores new technologies (e.g., FoundationDB, LSM‑trees, vectorized query engines), attends conferences, and shares knowledge via blogs or talks.

If you are ready to shape the backbone that powers observability for enterprises worldwide, join us at Datadog. Apply today and help us continue delivering the fastest, most reliable platform for developers, operators, and security teams.

Company Details

Company Overview

Employees: ████████
Founded: ████
Last Round: ████████
Amount: ████████
Company Insights · Gold

Upgrade to Gold to unlock company insights

About

Datadog is a worldwide leader in cloud‑native observability, empowering more than 6,500 customers—from startups to enterprises—to deliver reliable digital experiences. Founded in 2010, the company has grown to 6,500+ employees and a global footprint that includes offices in the United States, Europe, and Asia. Datadog’s flagship platform unifies infrastructure monitoring, application performance tracking, log management, security analytics, and AI‑driven telemetry into a single, scalable SaaS o ...
Post Date: September 24, 2025