Building production AI • Researching how it fails
MONISHA
KOLLIPARA
Mahindra University ’26 • HigherSelf • Salesforce ’26
About
My work sits at the intersection of LLM evaluation and reliable agentic systems. I build evaluation pipelines for tool-using LLMs, translate observed failures into measurable tests, and design multi-agent architectures that don’t quietly fail when it matters.
I’m Monisha Kollipara, a final-year AI undergrad at Mahindra University. Before HigherSelf, I built cross-framework analytics platforms at Salesforce and interned at Hike on data engineering.
Won ten hackathons across national and international stages. The biggest was the Sopra Steria International Student Challenge in Madrid, where our TerraTide project was called “a concrete solution to face the huge challenge of heatwaves” by jury member Yves Nicolas, CTO at Sopra Steria.
I also hold an Advanced Diploma in Fine Arts with distinction. Because not everything has to compile.
Fine arts taught me to see what’s actually there, not what I expect to find. The best evaluation work I’ve done started with the same instinct.
What I Work On
I study how AI fails
so it doesn’t when it matters.
I build evaluation harnesses that catch LLM failures before they ship. The failures that get through anyway, I write about.
Most of what I do sits at one question: how do you know your system is wrong before someone trusts it?
At HigherSelf, where I lead AI for two consumer apps, that question lives inside a generation pipeline that has to honour a deterministic foundation it isn’t allowed to contradict. At TerraTide, a multi-agent RAG system telling users what to do with their bodies in extreme heat, the question is retrieval grounding. A bad retrieval there is harm, not a UX issue. At claims_env, an evaluation benchmark I designed for agentic systems, the question is procedural compliance. Skipping a step is a failure that doesn’t announce itself.
Different domains, same question. That question is the whole portfolio.
Confidently wrong answers
The hardest failures aren't the ones that look broken. They're the ones that read fluent, pass surface checks, and contradict the underlying ground truth without flinching. At HigherSelf, most of my system is built to catch this: multi-pass generation against a deterministic foundation, automated rewrite triggers when the output drifts from the source it's supposed to honour.
Hallucination-to-action risk
When the user can't fact-check the system, hallucinations don't stop at the screen. They become actions. TerraTide tells people what to do with their bodies in extreme heat. In TerraTide, bad retrieval isn't a UX problem. It's harm. I designed it to force grounded recommendations, schema-validated tool calls, and failures that surface visibly rather than get buried.
Silent procedural failure
Most LLM benchmarks test whether the model knows the answer. Real agents are supposed to follow procedures, and skipping a step is a failure that doesn't announce itself. claims_env is an evaluation environment built around this. Agents have to follow real procedures (read policies, verify eligibility, detect fraud), and they fail in ways conventional eval doesn't catch.
Accepted Work
An AI-Driven Solution to Strengthen Heat Adaptation
AI-Enabled Risk Communication for Urban Heat Adaptation: Integrating Personalization and Citizen Participation
Experience
Where I’ve built things
Lead AI Engineer
HigherSelf
- Building production LLM systems for two consumer AI apps: Aatman (live on App Store) and Anantya (releasing shortly). Multi-pass generation constrained against deterministic ground truth. Every claim has to trace back to a computed fact. No hallucination by architecture, not just by prompting.
- Built a multi-layer post-generation guardrail stack: style enforcement, deterministic critic with structural checks, privacy and ethics gates, and global scrubbers. Critic triggers targeted rewrites on quality failures without full regeneration.
- End-to-end structured telemetry (latency, token counts, cost, cache hit rates, quality-gate outcomes) used to audit the pipeline, find bottlenecks, and drive a two-stage rendering architecture that took perceived first-content latency from ~70s to under 1s.
AMTS
Salesforce
- Returning full-time. Previously interned Summer 2025.
AMTS Intern
Salesforce
- Built a universal analytics embedding platform that runs across React, Angular, and Vue from a single codebase. Runtime framework switching. Config-driven dashboard embeds. Shipped CI-ready, deployed on Heroku.
- Designed a JSON-configurable embedding layer (auth, visual filters, bulk asset config, export workflows). Took new-embed integration from ~10 hours to ~10 minutes.
- Created automated test-generation agents as an evaluation harness for cross-framework parity. 120-250 Jest/Jasmine/Vitest tests per release cycle. Behavioral deltas surfaced early, regressions caught before they shipped.
- Selected for Worldwide AI + Tableau Day (global showcase). Presented the project and outcomes to a broader internal audience.
Data Engineering Intern
Hike
- Built the data governance layer for Hike's BigQuery infrastructure.
- Automated metadata generation, data discovery pipelines, universal Dataplex cataloging, Google SSO for unified access control.
- Reduced query costs and manual overhead while improving data security and visibility across the platform.
Projects
Things I’ve shipped
TerraTide
How do you ground health guidance when the retrieval source is noisy?
A multi-agent RAG system (retrieval agent, grounding critic, alert composer) that converts real-time weather signals and public-health data into personalized heat-risk guidance. Designed for safety under low-attention conditions: automated monitoring triggers alerts via email/SMS/voice, reducing hallucination-to-action risk by forcing grounded recommendations.
Accepted Poster, AI+Environment Summit, Zurich 2025
Accepted Abstract, INTROMET 2025, IITM Pune
Anantya & Aatman
What does reliable multi-pass LLM generation look like in production?
Two consumer AI apps where the generative layer isn't allowed to contradict a deterministic foundation. Most of the engineering effort is evaluation and quality assurance, not generation. The system catches its own mistakes before users see them. Aatman is live on the App Store; Anantya is releasing shortly.
NarrativeAI
Can a multi-agent pipeline turn raw news into evolving intelligence?
The hard problem was persistence: how do seven agents (Ingestion, Entity, Synthesis, Archetype, Contrarian, Ripple) maintain a shared, evolving understanding of a story across time? The system builds living dossiers that remember context, track entity relationships via D3.js graphs, and surface contradictions across 8 languages.
claims_env
How do you benchmark an agent that has to follow procedures, not just answer questions?
Claude Sonnet 4 scores 0.789 on this benchmark. That number is the point: most LLM evals test knowledge recall, but real-world agents need to follow multi-step procedures (read a policy, verify eligibility, detect fraud), and the failures don't announce themselves. 9 tasks, procedural generation, difficulty-adjusted scoring, 83 unit tests.
TheSecondMind
Can agents do iterative research with memory and self-critique?
The core question was whether agents could improve their own research output through structured self-review. Each cycle generates hypotheses, pulls from external sources (arXiv, NASA, Google), then a separate agent critiques and scores the output. Only results that pass a quality threshold survive to the next iteration. The interesting finding: meta-review feedback consistently improved output quality, but only when the critic agent had access to the original query context.
Recognitions
Selected wins and fellowships
Winner, StudentTribe Codeathon
2025First prize. Received on-the-spot SWE Intern offer at Raviga LLC.
Runner-up, Sopra Steria International Challenge
2025Second place globally in an international hackathon on "AI for a Sustainable Future."
Runner-up, "Building Agents that Learn Like Humans"
2025The Second MIND hackathon by Bosch and IIT Hyderabad. Received an AI Research Internship offer.
Runner-up, WATCH Hackathon
2024US Consulates, CUNY Bronx, City University of New York. Heatwaves and sustainable cities.
DESIS Ascend Fellow, D.E. Shaw India
2023Selected as 1 of 80 from 17,700+ applicants. Six-month mentorship program.
ICPC AlgoQueen, Rank 48 Asia
2023The Girls Programming Cup. Pan-Asia competitive programming.
Merit Scholarship, Mahindra University
2024$4,438 awarded for outstanding academic performance.
Second Position, Hult Prize
2023CEO of TwiceStyled, a startup pioneering sustainable fashion by transforming used denim into unique accessories.