Job Description: Elasticsearch & Observability Platform Architect
Role Overview
We are seeking an experienced Elasticsearch & Observability Platform
Architect to design and build a custom, large-scale observability
platform using Elasticsearch and OpenTelemetry standards. The role
focuses on building high-throughput ingestion pipelines, designing scalable
index/storage strategies, and optimizing query performance for massive
volumes of logs, metrics, and traces.
This is a hands-on architecture role—ideal for someone who understands
distributed systems deeply and can create a robust, scalable observability
backbone.
Key Responsibilities
1. Observability Platform Architecture
Architect the end-to-end observability ecosystem using:
o Elasticsearch (search, analytics, storage)
o OpenTelemetry (OTLP logs, metrics, traces)
o OTel Collectors, Kafka/streaming pipelines, or other ETL components
Define standard observability schemas, OTel semantic conventions,
and normalization rules.
Build multi-source ingestion pipelines from applications, microservices,
Kubernetes, and infra.
2. Elasticsearch Index & Data Architecture
Design index strategies for very large data volumes (billions of
documents/day):
o Index templates & mappings
o Shard sizing & allocation
o Hot-warm-cold (or frozen) storage architectures
Build ILM policies for retention, rollover, TTL, and tiering.
Implement compression, downsampling, and cost-optimized storage
patterns.
3. High-Performance Ingestion Architecture
Design ingestion pipelines for high-throughput log/metric/trace data:
o OpenTelemetry Collectors (processors, pipelines, exporters)
o Kafka, Logstash, FluentBit, Vector (as required)
Establish ingestion backpressure handling, buffering, and throttling
rules.
Ensure data quality: enrichment, correlation, normalization,
deduplication.
4. Query & Analytics Optimization
Design query models for large-scale observability datasets:
o Efficient use of keywords vs text fields
o Aggregation-heavy queries optimization
o Cardinality reduction techniques
Build query templates for:
o Log analytics
o Trace root-cause exploration
o Metric dashboards
o Multi-tenant or domain-based access
Optimize cluster performance for:
o High concurrency dashboards
o Low latency search/aggregations
o Large time-range analytics
5. Reliability & Operations
Perform cluster sizing, performance benchmarking, and scalability
modeling.
Troubleshoot ingestion bottlenecks, query timeouts, shard imbalances,
and memory pressure.
Define monitoring and health-check strategies for ingestion pipelines
and Elasticsearch clusters.
Required Skills & Experience
Core Expertise
7+ years with Elasticsearch, including:
o Cluster design
o Index & mapping architecture
o Query DSL + aggregation tuning
o Hot-warm-cold tiering
Strong expertise with OpenTelemetry:
o OTel Collector processors, pipelines, exporters
o OTLP ingestion for logs, metrics, traces
Experience handling large-scale observability data (hundreds of GB
to multiple TB per day).
Data Pipeline & Systems
Strong experience with:
o Kafka (preferred)
o Logstash / FluentBit / Vector
o Kubernetes-based deployments (ECK is a plus)
Solid distributed systems foundations: storage, memory, network
tuning.
Short Posting Summary (for LinkedIn/Portals)
We are hiring an Elasticsearch & Observability Platform Architect to
build a custom in-house observability platform using Elasticsearch
and OpenTelemetry. The role involves designing scalable ingestion pipelines,
advanced index/mapping strategies, and high-performance query models for
massive observability datasets. Experience in Elasticsearch
internals, OTel collectors, and large-scale distributed systems is required.