AI that shows its work

Rohit Agrawal — principal-level engineer, fourteen years across distributed systems and production reliability, now building independent AI systems that cite their sources, disclose their uncertainty, refuse what they cannot support, and gate their own releases.

Experience
14+ yrs — Oracle, Amazon, LimeRoad, Mobileum, Snapdeal, Subex
Most recent
Oracle Principal MTS · 2019–2026
Since April 2026
Building independent AI systems, full time
Seeking
Senior / Principal — AI Platform · LLMOps · Forward-Deployed AI · AI Quality Engineering Architect · AI Test Automation & Agentic AI Leader
Base
Bengaluru, India · IST (UTC+5:30)
A full-length tailored figure assembled from the portfolio itself: each garment is one system, named in the strip below.

CiteVyn

“Can I trust this answer — and trace every claim to its source?”

Citation-grounded Q&A over official AI documentation. Answers quote their sources verbatim; where no source supports an answer, CiteVyn refuses instead of guessing. Index updates reach production only through an evaluation gate.

A document figure: a quoted answer with numbered citation tabs, a refusal slot left empty where no source exists, and a 50-of-50 golden release gate at its foot. 1 2 3 REFUSED — NO SOURCE returned instead of guessed GOLDEN GATE 50 / 50 — a red case blocks the ship
State Live — cold-starts
Answers Quoted verbatim
Citations Every claim
Refusal A feature
Release gate 50/50 golden
Tests 361 passing
Stack FastAPI · pgvector

Quorum‑AI

“What happens when four models disagree about your question?”

One question runs against four models in parallel; they critique one another for two rounds, and a synthesis returns consensus, disagreement, source support, uncertainty, and a recommendation. The cost is approved before anything runs, and any fallback or simulation is disclosed, never hidden.

Four model columns debate across crossing threads and converge into one synthesis block with five fields; a cost tag hangs above — approved before anything runs. COST SHOWN FIRST nothing runs unapproved GPT-4O-MINI HAIKU 4.5 GEMINI 2.5 DEEPSEEK 3.1 CONSENSUS DISAGREEMENT SOURCE SUPPORT UNCERTAINTY RECOMMENDATION
State Live
Models 4 in parallel
Debate 2 critique rounds
Verdict 5 fields
Cost Approved first
Memory None — ephemeral

SaafSaans

“Is it safe for me to go outside right now — and if not, when?”

A Delhi-NCR air-quality companion that scores your risk — age, condition, planned activity — rather than the city's average, across 21 stations, and answers questions with cited health guidance. Every mode is labelled: live, deterministic fallback, or sample. The Hindi draft ships behind a banner saying no Hindi speaker has reviewed it yet.

Entered at Build with AI (Elastic × GDG Cloud New Delhi, 18 July 2026) as a four-tab Streamlit app that already existed, then rebuilt over the next three days into what runs today. It lives on one small machine that scales to zero when idle. Right now that machine is not answering — the address resolves and the server accepts the connection, then closes it without sending anything. The fault is being investigated. This page says so rather than leaving you to find out by clicking.

A standing column of six air-quality bands from good to severe; a small human figure at its base carries a personal risk ring; a bracket marks the best window to go outside; a roll of 21 ticks marks the stations covered. GOOD SATISFACTORY MODERATE POOR VERY POOR SEVERE BEST WINDOW is forecast YOUR RISK — NOT THE CITY'S 21 STATIONS, WORST-FIRST
State Deployed
Risk Per person
Scale CPCB bands
Timing Best window
Coverage 21 stations
Hindi Gated — unreviewed

NarraTwin AI

“Can project knowledge become a walkthrough without inventing a claim?”

Grounded walkthrough generation with citations, claim evaluation, consent checks, and release gates that run before anything is generated. Its own release-readiness review currently reads No-Go — so it is not deployed, and this page says so.

It is shown anyway, because the gate holding is the point.

A ghosted film frame held shut by a gate with three latches — claim evaluation, consent, release readiness — stamped No-Go by its own review. CLAIM EVAL CONSENT RELEASE NO-GO HELD BY ITS OWN GATES — WHICH IS THE POINT
State Phase 1 — No-Go
Gates Claim · consent · release
Deployment None — local only
Distribution Withheld by design

Carried, not shown

Two systems are still being built. They stay closed until they can be judged on finished work rather than on intent. They are named because a record that discloses its gaps should also say what exists — and the question each one is built to answer costs nothing to state.

EvalAxis

“Why can a failing test stop a release, when a measured drop in answer quality cannot?”

Evaluates LLM, RAG, and agent changes with evidence — and blocks CI on a quality regression.

In progress · closed

Aegis Contracts

“What should one AI system be allowed to promise another — and who checks?”

Early-stage work on contract-shaped guarantees between AI systems.

In progress · closed
Two zipped garment bags on a rail, each carrying a name tag. EVALAXIS AEGIS-CONTRACTS

One message away

Open to senior and principal roles in Forward-Deployed AI, AI Platform Engineering, and LLMOps / AI Reliability — and closely aligned AI quality and platform work.

Based
Bengaluru, India · IST (UTC+5:30)
Open to
Global relocation · international travel