Blog Archive

Everyone Is Building an AI AUTOSAR Toolchain. But How Should We Evaluate One?

Generating plausible AUTOSAR artifacts is an interesting capability. It is not yet evidence of a production-ready engineering system. We need a way to measure the difference.

Hybrid Databases Were the Wrong Fix for Requirements RAG

A negative infrastructure result from a requirements RAG lab: Postgres, ParadeDB, and OpenSearch could host useful pieces, but none replaced the current retrieval stack without losing ranking quality or evidence bundles.

The RAG Verifier Learned to Say No. It Still Missed a Hard Requirement Failure.

A held-out requirements RAG experiment: claim-level verification reduced harmful published answers from 35.7% selective risk to 4.0%, but strict human audit still found semantic completeness failures.

Citations Were Not Enough for Safety-Critical Requirements RAG

A negative result from a requirements RAG lab: even with structured records, cited JSON answers, and deterministic guards, the answerer failed the abstention gate when evidence was insufficient.

I Flattened 13,244 Requirements Into Chunks. 97% Lost Their Meaning.

A practical retrieval study over structured automotive requirements: why flatten-and-chunk failed, where ordinary top-k search stopped being the right tool, and how query routing improved both quality and latency.

I Tested a Simple RAG Idea: One Summary per Document. It Was Useful, But Not Enough.

A practical retrieval experiment: can an LLM-generated document summary replace chunk-level search? On a 5,000-document corpus, the answer was measurable — useful signal, clear ceiling, and a better baseline.

My Two AI Agents Talk MCP to Each Other. There Is No Standalone Tool Server.

The default MCP story is one fat tool server, many clients. When I needed two of my domain agents to talk to each other, that shape was the wrong one — and the alternative shows what MCP is actually good for.

I Built an AI Agent for Myself. My Colleagues Wanted to Use It. That Is Where the Problems Started.

When a domain expert builds a working AI agent and people start depending on it, the bus factor problem arrives long before the platform — and a different kind of work begins.