An AI audit that tests the system, not just the paperwork
We run hundreds of conversations against your chatbot, your agent or your scoring system to see what it actually does: whether it treats two people differently when only their name changes, whether it makes up terms that don't exist, whether it says it is an AI. From what we find, we write a report that works for both your engineering team and your legal counsel.
A responsible AI audit checks how an artificial intelligence system behaves against the principles the company says it follows and against the obligations of the EU AI Act. At Soamee we do it as engineers: we inventory the company's AI systems, including those that come built into SaaS tools, classify them by risk level and run test batteries on the ones that matter. We measure bias with counterfactual tests, detect made-up answers, review the transparency that Article 50 requires from August 2026 and the human oversight that high-risk systems will need from December 2027. The output is a report with evidence and a prioritised action plan. The final legal assessment belongs to the company's lawyer.
Six things you can verify with tests
A statement of ethical principles takes an afternoon to write and guarantees nothing. We look at what the system does when someone uses it, and every finding comes with the conversation or log entry that proves it.
Inventory and risk level
We list every AI system in use: the ones you built, the model APIs you call and the AI features inside your CRM, ATS or helpdesk. For each one, whether you are the provider or the deployer, and which EU AI Act risk level it falls into.
Bias, with counterfactual tests
We send the same requests changing a single detail: name, gender, age or origin. If the answer, the tone or the decision changes, we measure it and document it with the pairs of conversations.
Made-up answers and promises
We look for claims the system makes without any basis: return policies that don't exist, deadlines, prices, diagnoses. If a chatbot promises a refund, the company is liable for that promise.
Transparency (Article 50)
Whether the chatbot says it is an AI before the user asks, whether generated content is labelled, and whether the notice holds up when someone tries to get the system to pretend to be a person.
Real human oversight
When a person approves what the AI proposes, we measure how long they take, how often they correct it and whether they have the information to do so. Approving 99% of cases in four seconds does not count as oversight.
Traceability and data
What gets logged for each decision, for how long and who can look it up. And which personal data goes out to the model provider, in which region it is processed and under what contract.
What the EU AI Act requires, and when
Dates after the Digital Omnibus approved in July 2026, which pushed back the high-risk rules but left transparency and the prohibitions unchanged.
| From | What applies | What we check |
|---|---|---|
| 2 Feb 2025 | Prohibited practices (Art. 5) and AI literacy (Art. 4) | That no system manipulates people, does social scoring or recognises emotions in the workplace. What training the people using AI have, and whether it is documented. |
| 2 Aug 2026 | Transparency (Art. 50): chatbots, synthetic content and deepfakes | That the system identifies itself as an AI and that generated content is labelled where it should be. |
| 2 Dec 2026 | Machine-readable marking for generative systems already on the market | That generated images, audio and video carry a technical mark, on top of the visible notice. |
| 2 Dec 2027 | Annex III high-risk: employment, credit, insurance, education, essential services | Risk management, logging, human oversight, data quality and a draft of the fundamental rights impact assessment (Art. 27). |
| 2 Aug 2028 | High-risk AI built into regulated products (Annex I) | Which part of the product's conformity assessment depends on the AI component. |
What you get
A document with six sections. Management reads the first two; the rest is for the team that has to fix whatever turns up.
The report is not a legal opinion or a certificate. For most systems, the EU AI Act does not provide for any certificate. It is the technical evidence of how your AI behaves, which is what your legal counsel needs to assess compliance.
- 01
Executive summary
One page: which systems you have, which ones have problems, what is urgent because of a deadline and what can wait.
- 02
System inventory
Each system with its provider, your role under the EU AI Act, its risk level and the date from which each obligation applies to it.
- 03
Test results
What we ran, how many times and how we evaluated it. Every failure comes with the conversation or log entry that proves it.
- 04
Obligations map
Per system: which article applies, what is already met and what is missing.
- 05
Action plan
Prioritised by risk and by date. It separates what gets fixed in code, what is a process change and what your lawyer needs to review.
- 06
Reusable test battery
The test cases stay with your team, so you can run them again every time you change model or prompt.
Engineering and ethics in the same team
The audit is led by Javier Manzano, founder of Soamee. He is a software engineer studying Philosophy, and that combination has a practical use: turning a principle like 'do not discriminate' or 'a person must be able to step in' into a test that runs, gets measured and can be repeated.
At Soamee we build AI agents, chatbots and language model pipelines for clients. The audit uses the same techniques we use in our own projects: test case batteries, evaluation with an LLM as a judge calibrated against human review, and a log of every answer.
We are not a law firm. We work with the company's legal counsel: we show what the system does, and they decide what it means legally.
From inventory to action plan
We start by finding out what AI the company runs. There is usually more than management thinks.
Inventory
Interviews with the teams and a review of the tools you pay for. The result is the list of systems, their risk level and which ones deserve testing.
Testing
We design the test batteries and run them in a test environment: bias, made-up answers, transparency and human oversight.
Report
We deliver the report and go through it with management, with the engineering team and, if you want, with your legal counsel.
Review
When you change model or prompt, the behaviour changes. We run the battery again and compare it with the previous run.
Frequently asked questions about AI audits
Is an AI audit mandatory? +
The EU AI Act does not require every system to be audited. It does require transparency for chatbots and generated content from August 2026, and a risk management system, logging and human oversight for high-risk systems from December 2027. The audit is how you find out whether you comply, and how you have written evidence if anyone asks.
Does it replace legal advice? +
No. We document how the system behaves and which technical obligations apply to it. The legal interpretation and the legal liability belong to the company's lawyer, and the report is written so they can work from facts.
Can you audit a SaaS tool we didn't build? +
Yes, from the outside: we test its behaviour with your configuration and your test data, and we ask the vendor for the documentation the EU AI Act requires them to give you. If you use an ATS that ranks candidates, the system is high-risk even though someone else built it, and you have obligations as the deployer.
Do you need access to real customer data? +
No. Whenever possible we work in a test environment with synthetic data. If real logs need reviewing, that happens under a confidentiality agreement and with personal data minimised.
What does ethics have to do with the EU AI Act? +
The EU AI Act turns part of the ethical principles the European Commission published in 2019 into obligations: human oversight, technical soundness, privacy, transparency, non-discrimination and accountability. The rest falls outside the law and depends on what the company decides. The audit covers both, and the report separates what is mandatory from what is recommended.
What happens if you find a serious problem? +
We tell you as soon as we see it, without waiting for the report. If the fix is technical, we can do it ourselves or write it up as a spec for your team.
You may also be interested in
Tell us which AI you use
Tell us which systems you have, even if the list is half done: the chatbot on your website, the AI feature in your CRM, the CV filter. We'll tell you which ones deserve testing and where we would start.
Tell us which AI you useTell us your challenge. We'll propose a solution.
No commitment. Within 24 hours, you'll receive a proposal with scope, timeline and budget. No fine print.