Your AI agents, well governed.
Proof, not promises.
We adversarially test your agents before production and hand you an evidence file your auditor can re-examine. An agent stuck in pilot? We unblock it. An agent still to build? We build it — governed from design.
When the board asks "can you defend this AI?", you answer yes.
A vendor that vouches for its own AI isn't producing proof — it's marketing.
Locked-in proof
First-party proof lives inside the vendor's dashboard — no outsider can re-examine it.
Governance theatre
Policies and spreadsheets that prove nothing about the agent's real behaviour. A written policy is not proof.
Blind deployment
Shipping an agent to production untested — then discovering the problem after the incident.
This is no longer a matter of principle. It is a date.
The Autorité des marchés financiers' expectations on artificial intelligence governance take effect for the institutions it supervises.
— months left
Guideline E-23 on model risk management takes effect — AI models are explicitly in scope.
— months left
An agent already running on the day the rules take effect will have to be documented retroactively, on behaviour that was never measured. Testing before production costs a fraction of reconstructing it afterwards. Dates to be confirmed against the text in force at the time of your engagement.
An agentic AI governance firm.
A vendor's self-attestation isn't proof. We produce evidence that can be checked — independent of the vendor, re-examinable by your auditor.
Evidence, not our word
The file's value doesn't rest on our word: it's built to be re-examined piece by piece and re-run by a qualified third party on your side. We're not the arbiter — we produce the evidence.
Tested before, not watched after
We stress-test the agent adversarially before production — multi-turn attacks, traps, decision gates. Defensible before the incident, not documented after.
A firm, not a platform
Not a twelfth dashboard to run yourself. A firm that takes a position, signs a methodology and delivers a file — comparable, legible, documented.
One line in your code. Your agent becomes auditable.
Keep your stack. Our instrumentation is passive and framework-agnostic: once it's in, everything the agent decides and does becomes a re-examinable trace.
- Framework-agnostic — your agent code stays yours.
- Passive — no change to business logic.
- Runs on your side — your data doesn't leave.
- CI-ready — a decision gate on every release.
# Votre agent, tel quel from votre_app import agent # + une ligne from aurelia import audit agent = audit(agent) # ← tout ce qu'il fait # devient une preuve agent.run(task) # trace ré-examinable, # prête pour le dossier
Representative integration — passive, framework-agnostic, runs on your side.
From adversarial attack to file.
We don't grade isolated outputs: we pressure the agent across its whole trajectory, then link every finding to a trace.
Multi-turn attacks
Realistic, documented and replayable attacks put the agent to the test: hostile conversations, traps, composite attack chains.
Decision gates
Every evaluation resolves to a clear gate — pass, review or block — before any go-live.
Findings linked to traces
Every finding attaches to a timestamped trace: the file gets re-examined, not taken on trust.
Our verification layer builds on open-source, battle-tested and replayable methods — a deliberate method choice: neutrality, no lock-in, re-runnable by a third party.
The Agent Deployment Proof File.
A productised, repeatable deliverable, built to hold up before a regulator, a board or a risk committee — and to be re-examined piece by piece by your own auditor.
Does not constitute an assurance or audit engagement in the regulatory sense.
- Findings report linked to timestamped traces
- Documented replay set (attacks, configurations)
- Re-verification protocol for your auditor
- Findings mapped to applicable frameworks
- Re-verification workshop with your teams
Every finding is linked to the text that requires it.
The file does not read like a tool report. Every observed behaviour is tied back to the framework that applies to you — so your auditor works in their own vocabulary, not ours.
A mapping between observed behaviour and the status of a control. This is not an attestation of legal compliance — no tool can produce one, and we do not.
Defined by the situation, not the sector.
Any organisation where an AI agent is already in production or pilot, touching money, sensitive data or consequential decisions — and where someone must be able to defend it.
Agent doesn't exist yet? We build it — governed from design.
Custom AI agent development: scoping, guardrails, permissions, integrations, traceability. An integrator ships an agent that works in a demo; we deliver an agent that makes it to production — evidence file included.
We've built — and governed — our own agents.
DocFlow is a tool we designed, built and put into production for Quebec's notarial profession — a field where rigor and traceability are not optional.
We know where agents break because we've built them. It's our proof of competence — not a client, not a demo that "our governance works".
docflow.ca
- Document automation — Quebec notarial sector
- In real-world use, on real files
- Designed, built and governed by us
We train your teams to govern agentic AI.
Beyond the file: we make your teams autonomous. Understand the risks specific to agents that act — and learn to govern them.
AI governance fundamentals
The frameworks (ISO 42001, NIST AI RMF, EU AI Act, AMF/OSFI), roles and responsibilities, and what separates defensible evidence from policy on paper.
Agentic system risks
What changes when an agent acts: tool use (MCP), privilege escalation, agent-to-agent coordination, agent identity, and red-teaming.
Test and defend your agents
The method in practice: adversarial evaluation, decision gates, building a proof file, and engaging an auditor or regulator.
A productised deliverable, not vague hours.
Scoping
Agent scope, attack surfaces, applicable frameworks, tooling chosen with your auditor, evaluation plan.
Proof File
Full pre-deployment adversarial evaluation + re-examinable file mapped to frameworks + re-verification workshop.
Proof Keeping
Every change of agent, model or tool invalidates the proof. Re-testing on each material change, and periodically.
An agent still to build? Governed-from-design agent development sits alongside these tiers — evidence file always included.
Every engagement is scoped around a defensible, productised deliverable — never vague hours, never a guarantee.
What people ask us before we start.
What is AI agent governance?
It is the set of controls, guardrails and evidence that lets an organisation answer "yes" when asked whether it has a grip on an AI that acts: an agent that calls tools, triggers transactions or decides without a human validating every step. Governing an agent is not just an AI usage policy: it has to cover the agent's privileges, its tools, human oversight, traceability, and the evidence that it was put to the test before going into production.
What is an agent deployment proof file?
It's our core deliverable: a documented file gathering the agent's scope, the adversarial attacks we ran against it, execution traces, findings linked to those traces, decision gates (pass / review / block) and the mapping to applicable frameworks. The key point: it is built to be re-examined by your own auditor, not merely read. Reproducibility is partial and stated as such — we document what can be replayed and what cannot.
Do you certify or guarantee my AI agent's compliance?
No, and that's deliberate. Nobody can guarantee the behaviour of a probabilistic system. What we produce is defensibility: evidence that lets you justify your decisions to a board, an auditor or a regulator. Our work does not constitute an assurance engagement or an audit within the meaning of professional standards, and we are not a certification body — independence comes from your organisation's own auditor, who can re-examine the file.
Which frameworks do you work from?
Depending on your sector and exposure: the AMF's expectations in Québec, OSFI Guideline E-23 on model risk management, ISO/IEC 42001 and 42005, the OWASP Agentic AI Top 10, the NIST AI Risk Management Framework, and the EU AI Act where it applies. We don't sell a framework: we establish with you which ones apply, then map every finding in the file to the corresponding requirements.
When should you bring us in?
Before production, while the agent is still in pilot — that's when findings are cheapest to fix. We also step in on agents already in production when a regulatory deadline, a risk committee or an auditor calls for evidence — and at every material change afterwards: swapping the model, the tooling or the scope invalidates the previous evidence. And if the agent doesn't exist yet, we can build it — governed from the scoping stage.
Does this apply to my sector?
Our criterion isn't the sector, it's the situation: an organisation that already has an AI agent in pilot or production, touching money, sensitive data or consequential decisions, where somebody — risk, compliance, the board, a regulator, a professional order — has to answer for it. In practice that covers financial services and insurance, healthcare, professional services, the public sector and professional orders.
Where are you based, and do you work in French?
Aurelia Ops is a Canadian firm based in Montréal, Québec. We work in French and in English, on site or remotely, anywhere in Canada. Proof files, workshops and training are delivered in your organisation's language.
Deploy your AI agents — and stay defensible.
We'll pick an agent already in pilot or production with you, and show you what a proof file your auditor can re-examine looks like.
contact@aureliaops.comAI, well governed.
See what a proof file contains
Piece by piece: what your auditor receives, what they can re-run, and what we do not promise.