#Engineering Assurance
The Requirements RAG Router Failed. Then the Benchmark Failed Too.
A requirements RAG reliability experiment: heuristic routing failed on independent paraphrases, query grounding exposed a broken benchmark taxonomy, and typed evidence plans produced a narrow but useful result.
Read Post
Everyone Is Building an AI AUTOSAR Toolchain. But How Should We Evaluate One?
Generating plausible AUTOSAR artifacts is an interesting capability. It is not yet evidence of a production-ready engineering system. We need a way to measure the difference.
Read Post