crish nagarkar say hi

Hiwaving handI'm Crish.

ml engineer

Most of the effort in AI goes into making a model sound right. I work on the other half: finding out whether it actually is, and building the instruments that tell you when it isn't.

For the last two years that has meant , a legal AI that reads UK family-law judgments so solicitors don't have to. When I arrived it was citing the wrong authority with total confidence. The cause turned out to be . Rebuilding retrieval around that took correct-authority answers .

Before that I asked a smaller, more uncomfortable question: when we grade a model's reasoning automatically, does the grade mean anything? And if the judge can be wrong, the natural next question is whether the judge can be fooled — so my newest project : the scheming detector from a published safety eval, scored against mechanical ground truth instead of human labels.

The rest of the toolbox is open source. watches a running retrieval system and says when it has quietly gone bad. asks whether a model can write code a proof checker will actually accept, rather than code that merely reads as correct.

If there's a pattern, it's that I'm interested in the moment a system admits it doesn't know.

Work

Can LLM Reasoning Be Trusted? 2026 First-author paper on what automatic grading of reasoning is actually worth arxiv → AILES 2025– Production legal RAG over UK case law, built for solicitors who have to be right ailestech.ai → scorer-integrity 2026 Can the scheming detector be trusted? Auditing an AI-safety judge against mechanical ground truth — two bugs reported upstream, the whole campaign ran to $11 github → ragdrift 2026 Catches a retrieval system going bad before your users are the ones who notice github → dafny-eval 2026 Can a model write code a proof checker accepts? Mostly not, and the failures are legible github →

Experience

ML Engineer (Research Assistant in LLMs)

university of leeds · ailes

Owned retrieval and evaluation for a solicitor-facing legal AI: rebuilt hybrid search over 350,000+ chunks of case law, added calibrated abstention and token-level attribution so every conclusion cites its evidence, and distilled 300+ production failures into a fine-tuning dataset locked in by regression tests.

ailestech.ai — structured case intake
demo data, no real case visit the live product →

Store Retail Insight Analyst

primark · leeds

Python capacity models linking delivery schedules to physical backroom constraints cut intra-day stockouts 15% across a £50k+/week Click & Collect pipeline; a Power BI dashboard let management pre-position staff around predicted bottlenecks.

Machine Learning Engineer

miraiyantra · pune, india

Early-stage startup. Built the prototype ML pipeline for E-Vaidya, an AI-assisted wheelchair for cardiac patients — ECG/SpO₂ telemetry in, 1D-CNN anomaly classifiers out, Dockerised inference serving the live demo that won the £10,000 Young India prize.

Rooms

LegalTech in Leeds may 2026 One of seven invited to show AILES — the only university project among the startups press → AI SuperConnector 2024–25 Imperial College cohort of thirteen founders; £20k to try commercialising AILES showcase → Young India 2023 First prize for E-Vaidya — the wheelchair, the team, and a very heavy trophy

Off hours

Away from the terminal: badminton twice a week, a sketchbook that is mostly anime, and a cat who reviews all my work.

Badminton court mid-rally
tuesdays & fridays, leedssmash
meooow ♡ Tiku the cat, sitting upright and judging
tiku · chief reviewercute

Say hi

Readthe paper Runthe code Hireemail me