#Engineering Assurance

The Requirements RAG Router Failed. Then the Benchmark Failed Too.

A requirements RAG reliability experiment: heuristic routing failed on independent paraphrases, query grounding exposed a broken benchmark taxonomy, and typed evidence plans produced a narrow but useful result.

Everyone Is Building an AI AUTOSAR Toolchain. But How Should We Evaluate One?

Generating plausible AUTOSAR artifacts is an interesting capability. It is not yet evidence of a production-ready engineering system. We need a way to measure the difference.