Open to Director / VP Engineering conversations

Siva Samraj S

Director of Engineering at CultureMonkey. For 16+ years I have scaled engineering organizations and the platforms they run: 36+ engineers led across two companies, systems benchmarked to 20M+ concurrent connections and 120M+ events a day.

0+
Years
0+
Engineers led
0M+
Connections
0M+
Events / day
Siva Samraj S
Director of EngineeringCultureMonkey · Chennai · MIT Alumni
About

I scale engineering organizations
and the platforms they run.

For 16+ years I have built systems at the hard end of scale: 20M+ concurrent connections, 120M+ events a day, 35+ AI agents in production. Just as deliberately, I have built the teams behind them, growing a 16+ engineer organization at CultureMonkey today and a 20+ engineer organization at Sentienz Solutions before that.

My conviction is simple: architecture and organization are the same problem. Systems stay reliable when ownership is clear, teams ship fast when boundaries are well drawn, and both take deliberate engineering. That is the work I do.

Currently
Director of Engineering, CultureMonkey
Organization
16+ engineers across platform, product, data & AI
Platform
Global SaaS serving 10M+ enterprise employees
Scale journey · 2010 → today 2010201720212026 Sony · engineer to lead Sentienz · Founding Architect · 20+ engineers Akiro · 20M+ connections CultureMonkey · 16+ org, 10M+ users SCALE
Career Highlights

Numbers with the story behind them

Every metric here has an engineering decision, a team, and a production system behind it.

0+
Engineers led across two organizations
A 16+ engineer org at CultureMonkey today, and a 20+ engineer organization scaled at Sentienz as Founding Architect: hiring, mentorship, and delivery governance included.
0M+
Concurrent connections, benchmarked
The Akiro IoT platform, load-tested to 20M+ simultaneous connections with system-level TCP, kernel, and JVM tuning on AWS and Azure.
0M+
Events processed every day
Kafka and Spark pipelines moving 120M+ daily events into Elasticsearch, BigQuery, and S3 with zero data loss guarantees.
0+
AI agents running in production
Content generation, lead qualification, code review, and incident triage agents, built on RAG pipelines with multi-turn context.
0%
p99
Latency reduction at platform level
In-memory data grids, cache strategy, and p99-driven tuning turned slow request paths into real-time ones.
0%
Faster releases after re-architecture
Moving CultureMonkey from a monolith to event-driven microservices gave teams independent ownership and a 40% velocity gain.
Experience

16 years, three chapters

Mar 2025
— Present

Director of Engineering

CultureMonkey · Chennai, India
Scope: lead a 16+ engineer organization across platform, product, data, and AI. Own architecture and engineering OKRs for a global employee engagement SaaS serving 10M+ enterprise employees.
  • Re-architected a monolith into event-driven microservices, unlocking independent team ownership and a 40% increase in release velocity.
  • Built real-time feedback pipelines on Kafka, Redis, and ClickHouse, delivering 10x analytics performance with Elasticsearch ingesting 100K+ record batches.
  • Shipped AI-driven sentiment clustering across 5M+ feedback records and omnichannel survey delivery for Slack, Teams, WhatsApp, and Email.
  • Established full-stack observability with Datadog APM, Prometheus, and OpenTelemetry, plus mobile alerting for on-call incident response.
  • Run hiring, mentorship, and performance frameworks that grow senior engineers into leads, with OKRs tied to revenue, product, and reliability.
KafkaClickHouseRedisElasticsearchPostgreSQLMicroservicesAI AgentsDatadogOTEL
Oct 2017
— Feb 2025

Founding Architect & Sr. Engineering Manager

Sentienz Solutions · Bangalore, India
Scope: joined as Founding Architect and scaled the engineering organization to 20+ engineers, owning architecture, hiring, mentoring, and delivery governance across concurrent client platforms.
  • Architected the Akiro IoT platform, load-benchmarked at 20M+ concurrent connections through system-level TCP, kernel, and JVM tuning on AWS and Azure.
  • Built the Jarvis Central Data Platform: Spark and StreamSets pipelines feeding Elasticsearch, BigQuery, and S3 at 120M+ daily events with zero data loss.
  • Designed the OTTPlay analytics pipeline, from device telemetry through Kafka and Spark into Elasticsearch, so support teams query millions of events in real time.
  • Delivered RTRS, a real-time campaign engine with Google Maps geo-targeting, bidirectional notifications, and Apache Ignite in-memory compute for 100+ concurrent campaigns.
  • Stood up three-layer observability (Datadog, Prometheus, OpenTelemetry) and a bespoke mobile alerting system with severity triage and escalation chains.
KafkaSparkCassandraIgniteAerospikeRedisElasticsearchStreamSetsIoT / MQTTAWS · Azure
2010
— 2017

Software Engineer → Technical Lead

Sony · Bangalore & US onsite
Scope: grew from Software Engineer to Technical Lead over seven years, serving as the US onsite technical liaison for data platform modernization.
  • Led the Oracle to Hadoop Data Lake migration, cutting ETL processing time by 60% for enterprise BI workloads.
  • Introduced Kafka streaming pipelines that modernized batch-era BI into near-real-time reporting.
HadoopKafkaOracleMySQLBI Modernization
Signature Systems

Four platforms, told properly

Each one framed the way engineers actually evaluate work: the challenge, what was built, and what changed.

Akiro IoT Platform

IoT · Extreme Scale
Challenge
Multi-tenant IoT across healthcare, telematics, and energy, where every device must hold a live connection and no message can be dropped.
Built
An MQTT platform on Kafka, Cassandra, Ignite, and Aerospike with bidirectional notifications, tuned at the TCP, kernel, and JVM level.
Impact
Load-benchmarked to 20M+ concurrent connections with full p50/p95/p99 profiles on both AWS and Azure.
20M+connections120M+daily messages2clouds

AI Agents Ecosystem

Applied AI
Challenge
Repetitive knowledge work across marketing and engineering that scaled with headcount instead of software.
Built
A fleet of RAG-powered agents with multi-turn context and notification integration: content generation, lead qualification, code review, and incident triage.
Impact
35+ agents in production across two domains, quietly doing work that used to queue on people.
35+in productionRAGpowered2domains

OTTPlay Analytics Pipeline

Real-Time Analytics
Challenge
Support teams were debugging playback failures from raw logs, slowly, and only after users complained.
Built
Device telemetry flowing through Kafka, enriched in Spark with session stitching and error classification, indexed into Elasticsearch tuned for support query patterns.
Impact
Millions of events queryable by user, device, session, or error, giving instant root cause without touching a raw log.
01

Device Telemetry

App events, playback errors, buffering metrics emitted in real time.

02

Kafka Ingestion

Partitioned by user for ordered, zero-loss delivery.

03

Spark Processing

Session stitching and error classification in-stream.

04

Elasticsearch

Mappings and shards tuned to support query patterns.

05

Support Queries

Millions of events searchable by user, device, or error.

M+events indexedReal-timequeriesZeroraw log access

Real-Time Response System (RTRS)

Geo · Campaigns
Challenge
Location-based campaigns needed sub-second targeting decisions with proof of delivery, at fleet scale.
Built
Google Maps and Directions APIs for route-aware targeting, bidirectional delivery receipts, and Apache Ignite in-memory compute over a 20+ node Hadoop cluster.
Impact
100+ concurrent campaigns with a 75% latency reduction on the hot path, monitored end to end with Datadog and mobile alerting.
100+concurrent campaigns↓75%latencyBidir.receipts
How I Lead

Principles I actually run teams on

Architecture is org design

Service boundaries and team boundaries are drawn together. When ownership maps cleanly to systems, teams ship independently and incidents have obvious owners.

OKRs tied to the business

Engineering goals connect to revenue, product, and reliability outcomes. Every quarter my teams can say what their work changed for the company.

Grow leads, not dependencies

Hiring, mentorship, and performance frameworks are built to turn senior engineers into leaders. I have done this twice, at Sentienz and at CultureMonkey.

Reliability is a culture

SLOs, observability, and on-call alerting are designed into platforms from day one. Datadog, Prometheus, and OpenTelemetry are standard equipment.

Skills

The toolkit, grouped honestly

Distributed Systems & Data

  • Apache Kafka & Spark · streaming at 120M+ events/day
  • Cassandra, Ignite & Aerospike · low-latency stores
  • ClickHouse & Elasticsearch · analytics & search
  • Redis · sub-ms cache, pub/sub
  • IoT & MQTT · 20M+ connections benchmarked
  • StreamSets · pipeline orchestration

AI & Intelligent Systems

  • AI agent engineering · 35+ in production
  • RAG pipelines & vector search
  • LLM application development
  • Enterprise chatbot architecture
  • AI-driven analytics · sentiment at 5M+ scale

Databases & Platforms

  • PostgreSQL, Oracle & MySQL · schema & query tuning
  • AWS & Azure · multi-cloud production
  • Ruby on Rails & JVM services
  • Performance engineering · kernel, TCP, JVM, p99

Leadership & Operations

  • Org design & scaling · 36+ engineers led
  • Hiring, mentorship & performance frameworks
  • Engineering OKRs · tied to revenue & reliability
  • Observability & SRE · Datadog, Prometheus, OTEL
  • Incident management · mobile alerting, runbooks
How my stack composes · reference architecture
Sources
IoT · MQTT devicesWeb & mobile appsSlack · Teams · WhatsAppHRMS & SaaS APIs
Ingestion
Apache KafkaStreamSetsBidirectional notifications
Processing
Spark StreamingApache Ignite in-memoryAI agents · RAG
Storage
CassandraClickHouseElasticsearchRedis · AerospikePostgreSQL
Serving
Real-time analyticsSearch & support queriesCampaigns & chatbots
Observability spans every layer · Datadog APM · Prometheus SLOs · OpenTelemetry tracing · mobile alerting
Education

B.E. Computer Science · Madras Institute of Technology (MIT)

Anna University, Chennai · MIT Campus, Chrompet

Contact

Let's talk about scale.

I am open to Director and VP of Engineering conversations, platform architecture challenges, and collaboration in distributed systems and AI infrastructure. Email is the fastest way to reach me.

Chennai, India (IST) YouTube · Sentienz Solutions