3 September 2026 · Stela Suils Cuesta
Why my EU AI Act tool missed a prohibition, and what fixed it
A tester's question about a voice-interview agent exposed a retrieval bug: the Article 5 prohibitions were in the knowledge base and unreachable. Three lessons for any RAG system over regulation.
A tester asked my EU AI Act tool a question I had not anticipated:
“Our recruiting agent interviews candidates by voice and noted one had a ‘neutral tone’. Our client thinks there’s a legal issue.”
There is. Two, actually.
Article 5(1)(f) prohibits AI that infers emotions in the workplace. That is a full prohibition, not a high-risk obligation. And Article 3(39) defines emotion recognition as inference from biometric data, which is why the answer turns on whether “tone” comes from voice audio (biometric, so within the prohibition) or from text (outside that definition, though recruitment is still Annex III high-risk regardless).
My tool got that right. What it got wrong first was worse: it analysed high-risk classification and never checked the prohibition at all.
The cause was mundane
The Act puts all eight prohibited practices inside one paragraph, so my chunker produced a single 4,754-character block, the largest in the corpus. Both keyword and semantic search buried it under short, topically dense recitals. The prohibition was in the knowledge base and unreachable.
The fix was structural: one chunk per prohibition, per definition, per Annex III point. Then a rule that employment questions always retrieve the prohibition, because the legal order of analysis is fixed, and no ranking algorithm knows that.
Three things I would take to any RAG system over regulation
- Chunk by the document’s own structure, not by character count. Law is already structured. Using a generic splitter throws that away.
- Recitals explain, articles bind. If your system cites a recital as the source of a duty, it is wrong in a way that reads as confident.
- Encode the order of legal analysis in the pipeline. Prohibited-practice checks come before risk classification. Embedding similarity has no idea.
The uncomfortable part
My evaluation metrics scored that broken answer highly. Neighbouring text made it look well-grounded. I only found it because someone asked a real question.
That is why FairAudit now binds every citation to the passage it came from, runs deterministic checks before a model gets a say, and publishes its evaluation scores including the questions it does worst on. If you have a hard question about the Act, I would rather hear that than give you a demo.
Regulatory information, not legal advice.