#LLM
The RAG Verifier Learned to Say No. It Still Missed a Hard Requirement Failure.
A held-out requirements RAG experiment: claim-level verification reduced harmful published answers from 35.7% selective risk to 4.0%, but strict human audit still found semantic completeness failures.
Read Post
Citations Were Not Enough for Safety-Critical Requirements RAG
A negative result from a requirements RAG lab: even with structured records, cited JSON answers, and deterministic guards, the answerer failed the abstention gate when evidence was insufficient.
Read Post
I Tested a Simple RAG Idea: One Summary per Document. It Was Useful, But Not Enough.
A practical retrieval experiment: can an LLM-generated document summary replace chunk-level search? On a 5,000-document corpus, the answer was measurable — useful signal, clear ceiling, and a better baseline.
Read Post