varun.devOPEN TO WORK5+ YRSJAVA · PYTHONAWS · K8SGENAI · RAGRetrievalThroughputResilienceTerminalBuiltTrajectorySignalRésumé

● OPEN TO WORK · Jersey City, NJ · Software Engineer · Backend & Distributed Systems · AI

Varun Kammadanam.

I build the backend systems and AI pipelines that turn messy data into fast, reliable software — Java/Python services, Kafka event platforms on AWS & Kubernetes, and production RAG with LLM agents. 5+ years across healthcare and finance.

This page isn’t a brochure — it runs scaled-down versions of systems I’ve actually built. Run a query, turn the dial, break something.

0
operational friction · GenAI doc intelligence

At McKesson I architected a GenAI document-intelligence platform (OpenAI + LangGraph agents + RAG) that extracts and validates unstructured medical records. Automating that path cut operational friction ~40% on the healthcare remediation platform.

0
mean time to detect · JPMC

At J.P. Morgan Chase I built the Datadog dashboards and Prometheus alerting across security microservices, and wired LLM-driven anomaly detection into compliance monitoring — halving the mean time to detect anomalies.

0
events sustained · zero loss

My cloud-native event-processing platform runs partitioned Kafka topics with Java consumer groups on Kubernetes and AWS Lambda for bursts. Idempotent consumers and disciplined offset management sustain 10k+ events/sec with zero message loss.

0
test coverage baseline

I hold a 90%+ JUnit/Mockito coverage baseline on the services I own, with CI quality gates — part of treating reliability and observability as the definition of done, not an afterthought.

/ 01 — RETRIEVAL

Ask my retrieval engine.

The RAG pipeline behind my GenAI document work and my Semantic Retrieval Engine. Pick a mode, run a query, and read the retrieval scores and per-stage latency.

awaiting query · pick a mode and hit Run.

/ 02 — THROUGHPUT

Push the event platform.

My Kafka event-processing platform — 10k+ events/sec, zero message loss. Drag the ingress load and watch consumers autoscale while p99 holds.

INGRESS LOAD2.5k events/s
consumers
4
p99 latency
38ms
cost / hr
$1.44
message loss
0
NOMINAL · consumer group balanced · offsets committed

/ 03 — SELF-HEALING

Watch a pipeline fix itself.

My Self-Correcting Document Pipeline. Feed it a clean document, or a messy one and watch validation route it back for a self-correcting retry.

idle · the graph classifies → extracts → validates → summarizes, with a bounded retry.

/ 04 — RESILIENCE

Break the backend. Watch it survive.

A request travelling through a distributed backend I’d build. Send it, then inject a failure and watch retries, replica promotion and caching keep it alive.

cluster nominal · click Send request

/ 05 — TERMINAL

Query the stack.

A real command parser. Try skills · projects · benchmark · kubectl get pods · deploy. help lists everything.

varun@stack — zsh
type help · ↑/↓ history · tab to complete

/ 06 — BUILT

Systems I’ve shipped.

Two open-source builds you can read on GitHub, plus production and independent systems. Click any card to expand the full case study.

Open source · mcp-docqa-server

Semantic Retrieval Engine

RAG retrieval shipped as reusable MCP infrastructure — index a corpus once, query it from any AI client.

  • recall@5 ≥ 0.90 (CI-gated)
  • 2 storage backends, 1 test battery
  • runs keyless or scales to Postgres

Hybrid vector + BM25 fused with Reciprocal Rank Fusion; interchangeable SQLite / Postgres+pgvector stores; pluggable OpenAI or keyless embeddings; a recall@k eval harness gates CI.

PythonMCPpgvectorOpenAIDocker
View on GitHub ↗
Open source · langgraph-doc-pipeline

Self-Correcting Document Pipeline

Multi-agent document processing on LangGraph that repairs its own failures.

  • self-correcting retry recovers messy docs
  • pluggable rule-based (keyless) or OpenAI engine
  • batch metrics: review rate, recovery, completeness

classify → extract → validate → summarize with a confidence-weighted retry loop; low-confidence docs route to review instead of fabricating fields; every node is auditable.

PythonLangGraphOpenAIpytest
View on GitHub ↗
McKesson

GenAI Document Intelligence

LLM agents + RAG turning unstructured medical records into structured data.

  • −40% operational friction
  • +45% data-validation throughput
  • 90%+ test coverage baseline

OpenAI models and LangGraph agents over a RAG pipeline; parse → chunk → embed → semantic search → extraction; delivered via REST/gRPC on Kubernetes, monitored with Datadog + Prometheus.

OpenAILangGraphRAGpgvectorKubernetes
J.P. Morgan Chase

LLM Compliance Anomaly Detection

LLM anomaly detection inside security compliance monitoring.

  • −50% mean time to detect
  • −15% deployment errors
  • reduced manual access-log review

High-frequency Kafka audit streams into concurrent Java parsers; an LLM layer flags anomalous access patterns; low-latency gRPC services on tuned Oracle/SQL Server with Datadog alerting.

JavaKafkagRPCOracleDatadog
Independent build

AI-Powered Document Assistant

Semantic search + extraction with RAG and agent orchestration.

  • secure REST + gRPC with RBAC
  • Pinecone + pgvector retrieval
  • grounded answers over enterprise docs

Java (Spring Boot) owns a secure REST/gRPC surface with strict RBAC; a Python (FastAPI) service runs LangGraph agents over RAG; embeddings in Pinecone + pgvector; full-path observability.

JavaFastAPIOpenAILangGraphPinecone
Independent build

Cloud-Native Event Processing

Event-driven microservices at 10k+ events/sec with zero message loss.

  • 10k+ events/sec sustained
  • zero message loss under load
  • horizontal scale on Kubernetes

Partitioned Kafka topics; Java consumer groups on Kubernetes; AWS Lambda absorbs bursts; DynamoDB state; idempotent consumers and offset discipline give at-least-once with dedup.

JavaKafkaAWS LambdaDynamoDB

/ 07 — TRACK RECORD

Where I’ve shipped.

Three companies, one throughline: production backend and AI at healthcare and bank scale. The numbers below are what I shipped — not what I aspire to.

McKesson

Sep 2024 — Present · Texas
Senior Software Engineer · Cloud Data & Enterprise Platform Engineering
−40%
ops friction
45%
throughput
25%
cloud spend
90%+
coverage
  • Architected a GenAI document-intelligence platform — OpenAI LLMs, LangGraph agents and RAG over unstructured medical records.
  • Built high-throughput REST & gRPC microservices on Kubernetes with semantic vector search for healthcare data validation.
  • Engineered distributed Java (Spring Boot) & Python services on AWS Lambda; Redis multi-level caching cut DB reads 30%.
  • Lead technical design reviews and mentor a team of 4 engineers; reduced cloud spend 25% at sub-second p95 latency.

J.P. Morgan Chase

May 2022 — Sep 2024 · New York City
Software Engineer · Financial Platforms & Security Engineering
−50%
MTTD
−15%
deploy errors
2.4yr
bank scale
  • Programmed fault-tolerant Java microservices securing distributed financial access-management platforms.
  • Integrated LLM-driven anomaly detection into security compliance monitoring pipelines.
  • Built parsers with Java concurrency over high-frequency Kafka audit streams; tuned SQL across Oracle/SQL Server.
  • Designed CI/CD (Jenkins, GitHub Actions) with Docker on AWS; halved MTTD via Datadog + Prometheus.

OpenText

Jun 2020 — Dec 2021 · Hyderabad
Software Engineer · Analytics & Enterprise Application Development
28%
throughput
  • Refactored legacy enterprise workflows into scalable Spring Boot microservices.
  • Managed high-volume DB migrations with optimized SQL and Hibernate/Spring Data JPA; fed Tableau & Power BI.
  • Added automated regression suites and CI quality gates, cutting post-deployment defects.

/ 08 — BLUEPRINT

How it connects & scales.

A representative production shape from systems I’ve built — not any employer’s proprietary diagram.

Request path
Client — React / TypeScript
API Gateway — authn · RBAC
Backend — Spring Boot · gRPC
Messaging — Kafka · event-driven
Cache — Redis multi-level
Database — PostgreSQL · pgvector
Cloud — AWS · Kubernetes · Terraform
Observability — Datadog · Prometheus
Production engineering
90%+
test coverage
p95 <1s
API latency
−25%
cloud spend
−15%
deploy errors
Observability, testing and cost discipline are part of my definition of done.

/ 09 — SIGNAL

Open a channel.

Open to Software / Backend / Senior SWE and AI / Applied AI / Generative AI Engineer roles.