top of page

Inside the AI SDLC: what each mechanism should prove

Empty modern office with rows of desktop computers facing large windows and a glowing city skyline at dusk.

The companion argument to "governance is evidence, not assurance" is a practical one. Evidence has to come from somewhere. In a governed AI SDLC, the AI Software Development Lifecycle, it comes from five mechanisms, and each produces a specific record that someone outside the team can inspect. Two questions decide whether a mechanism is doing its job: what it should prove, and what an assessor looks for to confirm it.


The five are not theoretical. In vehicle manufacturing, a version of them has been a legal condition of market access since 2022, which is about as close to proof as this field offers that they can be built and run at scale.


The premise throughout is that governance belongs to the pipeline that builds and runs the AI, not to a document filed beside it. Fewer than 1% of organisations have fully operationalised responsible AI [1], and the distance between a written policy and a running system is what these five mechanisms exist to close. Each turns a governance intention into something the pipeline emits as it works.


The five mechanisms of the AI SDLC

Alignment with regulations and standards

This mechanism has to show that the system conforms to the rules that apply to it, that the conformance holds continuously, and that it can be demonstrated clause by clause. In enterprise terms the reference point is ISO/IEC 42001, the first international standard for an AI management system, published in December 2023 [2]. It sets requirements across its management-system clauses and a set of Annex A controls spanning AI policy, impact assessment, the lifecycle itself and data governance, with a Statement of Applicability recording which controls apply and how. The evidence a pipeline produces here is a body of tests, reviews and risk assessments, each mapped to the control or obligation it satisfies and kept current as the system changes.


An assessor works backwards from the requirement. For a given control, they look for the tests that cover it, their latest results, and an unbroken chain from a risk to the control that addresses it to the verification that confirms it. A control marked satisfied with nothing behind it reads as a gap.


The vehicle-cybersecurity regime shows the standard in force. Under UN Regulation 155 and ISO/SAE 21434, a manufacturer holds a threat analysis in which every scenario resolves to a requirement and a verification result, indexed to the regulation's clauses, and the type-approval audit walks that index. It is the discipline ISO/IEC 42001 describes, already operating under legal obligation.


Operational controls

The proof this mechanism carries is that the checks governing the work run inside the pipeline, at each gate, instead of forming a queue of approvals at the end. The evidence is the gate record: which check ran on which change, the decision it returned, and the approval that released the work. Run well, this is where governance starts to save time, because a control that executes automatically at commit beats a review scheduled for the following week. The reference model is the govern, map, measure and manage loop of the NIST AI Risk Management Framework, built into the pipeline as gates rather than run as a periodic exercise.


What an assessor examines is, for each release, the gate decisions and the evidence that the right path was taken, and for any regulated change, confirmation that its approval was in place before the work shipped.


In the automotive version, a change is classified the moment it is raised. A release touching the type-approval-relevant software identifier is held until its extension is approved, while a routine change proceeds under the recorded software-update process. The classification and the gate decision are both retained.


Visibility

Here the requirement is that an organisation can see the state of its own AI on demand, without having to reconstruct it after something has gone wrong. The practice that has grown up around this is the AI bill of materials: a structured, signed record tying together the software supply chain (a software bill of materials), the model (its card, its versions, its evaluation results) and the data (sources, licences, lineage). It answers what a black-box deployment cannot: what is running, on what data, in which version, and what has changed. It also lets drift or a retraining trace back to a specific dataset version, and that traceability is what the security mechanism depends on.


The assessor’s test is a plain one. A current inventory, and the ability to answer, for any component or vulnerability, which systems use it and which versions sit in production. The measure is whether that answer is a query or a project.


The automotive regime bakes this in: a software bill of materials generated and diffed against the previous release on every build, and a per-vehicle configuration record resolvable from the software identifier, so that a new vulnerability yields the affected vehicles, versions and markets in minutes.

Circuit board in a test jig with probes beside an oscilloscope in a lab, suggesting electronics diagnostics.

Security

This mechanism proves that the system has been tested against how it can be attacked and how it can fail, before exposure, and that the testing carries on once it is live. In enterprise AI the discipline is adversarial testing, or red-teaming: probing the system with the inputs a real adversary would use, from prompt injection to tool misuse, and repeating it continuously as the model's behaviour drifts and its failure surface moves. The stronger programmes run this inside the delivery pipeline, catching a regression in safety behaviour the way they would catch a regression in function, and they watch for behavioural and bias drift in production rather than trusting that validated behaviour holds. A sandbox earns its place by finding failure in rehearsal.


An assessor looks for the threat model and its history of change, the test campaigns with results traced to the requirements they cover, and evidence that the system was exercised in a controlled environment before it was exposed. A model that meets its first real adversary in production has been launched, not secured.


The automotive regime does this through a threat analysis maintained per change, penetration testing of interfaces and protocols, and virtual validation, credibility-qualified under the newer automated-driving rules, that absorbs the regression load before anything reaches a vehicle, with a verified-before-install step and a tested rollback at deployment.

Close-up of hands writing on blank paper at a desk in a bright office, with a laptop blurred in the background.

Accountability

The proof here is that a person, not the system, answers for every AI-assisted action, and that the answer is on record before the action takes effect. This is the mechanism that keeps a human in charge. AI can accelerate the evidence, drafting tests, proposing mappings, generating documentation, but decision authority stays with people, and every AI-assisted output is attributable, reviewable and signed before it enters the record. The supporting practices come straight from software supply-chain security, applied to models: signing artefacts and manifests, publishing checksums to an append-only log so unauthorised change shows up, and treating the human approval as a recorded step.


For any decision, an assessor should be able to retrieve who signed it off and what informed it. Accountability is what allows the other four mechanisms to be trusted, since a pipeline can generate a great deal of evidence that no one, in the end, stands behind.

The automotive regime records it per vehicle: threat-assessment deltas and impact classifications are AI-drafted but assessor-signed before any build, and a ledger holds the system, the versions before and after, the software identifier, the timestamp and the outcome. It is the audit trail from a single change out to every vehicle in the field, with a qualified person behind each step.

The standards behind this TTMI's reference guides on the regulations discussed here, UN R155 and R156 and ISO/SAE 21434.

Why the AI SDLC only works as a pipeline

The five mechanisms are set out separately for clarity, but their value is in how they connect. Security reads from visibility. Accountability is what makes the rest worth trusting. Alignment and operational controls are one discipline seen from two angles, the rulebook and the workflow. Kept as five separate initiatives, they generate five kinds of paperwork. Held together in one pipeline, they produce a system that can account for itself.


That is also why so few organisations have got there. The discipline cannot be bought as a document or bolted on as a final step; it has to be built into how the work is done. The industries that adopted it under legal compulsion are the proof that it is achievable, and the clearest picture available of what good looks like.


TTMI's wider material on the discipline, from the mechanisms to the standards behind them, sits on the AI SDLC resource library.


Sources

[1] WEF AI Governance Alliance and Accenture, Advancing Responsible AI Innovation: A Playbook, 2025.

[2] ISO/IEC 42001:2023, Information technology, Artificial intelligence, Management system.

[3] UNECE, UN Regulations 155 and 156; UN Regulation 185 and GTR No. 26 (automated driving), adopted 2026.

 
 
 

Comments


bottom of page