From models to operating paths
Retrieval, evaluation loops, serving reliability, and the data quality that decides whether an AI product actually works.
I'm Raja Mannem — a Senior Systems Engineer for BigData & AI/ML at Amazon Web Services. Over 11+ years I've helped enterprises run large-scale data systems on AWS and was recognized as an Amazon EMR Subject Matter Expert. Now I build the AI layer on top: retrieval, evaluation, agents, and the pipelines that keep them fed. Built at Scale is where I write it down.
The through-line is turning messy operational constraints into systems people can reason about — whether the payload is a nightly batch job or a model in the serving path.
Retrieval, evaluation loops, serving reliability, and the data quality that decides whether an AI product actually works.
Lakes and lakehouses, streams, transformations, contracts, and lineage — the ergonomics that make data usable for ML.
EMR, Spark, Hive, Tez, HBase, Presto, plus Kinesis, MSK, DynamoDB, Glue, and Athena — and the edge cases that only show up at scale.
A decade-plus on the AWS Big Data front line — escalations and operations on large workload clusters for enterprise customers, with open-source fixes shipped along the way.
The clearest picture of how I think is in the notes, the projects, and the reading list. To get in touch, find me here: