AI that shows its work
Rohit Agrawal — principal-level engineer, fourteen years across distributed systems and production reliability, now building independent AI systems that cite their sources, disclose their uncertainty, refuse what they cannot support, and gate their own releases.
- Experience
- 14+ yrs — Oracle, Amazon, LimeRoad, Mobileum, Snapdeal, Subex
- Most recent
- Oracle Principal MTS · 2019–2026
- Since April 2026
- Building independent AI systems, full time
- Seeking
- Senior / Principal — AI Platform · LLMOps · Forward-Deployed AI · AI Quality Engineering Architect · AI Test Automation & Agentic AI Leader
- Base
- Bengaluru, India · IST (UTC+5:30)
- Relocation
- Open worldwide
CiteVyn, golden case claude_api_006
Refused
“What is the capital of France?”
- citations
- 0
- domain
- unsupported
- confidence
- none
- answered
- No — by policy
CiteVyn
“Can I trust this answer — and trace every claim to its source?”
Citation-grounded Q&A over official AI documentation. Answers quote their sources verbatim; where no source supports an answer, CiteVyn refuses instead of guessing. Index updates reach production only through an evaluation gate.
Golden regression run
52 / 52 passed
A red case blocks the index from reaching production.
- cases
- 52
- failed
- 0
- run
- 17 Jul 2026
- gate
- Blocks promotion
Quorum‑AI
“What happens when four models disagree about your question?”
One question runs against four models in parallel; they critique one another for two rounds, and a synthesis returns consensus, disagreement, source support, uncertainty, and a recommendation. The cost is approved before anything runs, and any fallback or simulation is disclosed, never hidden.
Production readiness review
Go
The review before it said No-Go. That one is kept too.
- decision
- Go — single instance
- dated
- 21 Jun 2026
- scope
- MVP, small user base
- superseded
- No-Go of 16 Jun
SaafSaans
“Is it safe for me to go outside right now — and if not, when?”
A Delhi-NCR air-quality companion that scores your risk — age, condition, planned activity — rather than the city's average, across 21 stations, and answers questions with cited health guidance. Every mode is labelled: live, deterministic fallback, or sample. The Hindi draft ships behind a banner saying no Hindi speaker has reviewed it yet.
Entered at Build with AI (Elastic × GDG Cloud New Delhi, 18 July 2026) as a four-tab Streamlit app that already existed, then rebuilt over the next three days into what runs today. It lives on one small machine that scales to zero when idle. Right now that machine is not answering — the address resolves and the server accepts the connection, then closes it without sending anything. The fault is being investigated. This page says so rather than leaving you to find out by clicking.
Measured at head, today
Counted, not claimed
Every number here has a command that checks it.
- test functions
- 628
- test files
- 32
- seeded advisories
- 43
- commits
- 161
NarraTwin AI
“Can project knowledge become a walkthrough without inventing a claim?”
Grounded walkthrough generation with citations, claim evaluation, consent checks, and release gates that run before anything is generated. Its own release-readiness review currently reads No-Go — so it is not deployed, and this page says so.
It is shown anyway, because the gate holding is the point.
Release readiness review
No-Go
No-Go for production release. No release tag has been created.
- dated
- 1 Jul 2026
- release tags
- 0
- blocked
- Paid providers, video export
- allowed
- Local mock demo
Carried, not shown
Two systems are still being built. They stay closed until they can be judged on finished work rather than on intent. They are named because a record that discloses its gaps should also say what exists — and the question each one is built to answer costs nothing to state.
EvalAxis
“Why can a failing test stop a release, when a measured drop in answer quality cannot?”
Evaluates LLM, RAG, and agent changes with evidence — and blocks CI on a quality regression.
In progress · closedAegis Contracts
“What should one AI system be allowed to promise another — and who checks?”
Early-stage work on contract-shaped guarantees between AI systems.
In progress · closed