Agentic AI Governance in GxP Life Sciences: The Reasoning-Layer Control Gap

Why validated pharmaceutical and biotechnology workflows can still act on an unauthorized interpretation, and what has to be checked before execution.
Classifying a biologics deviation as major can send the case down a very different path from classifying it as critical. A critical finding may trigger quarantine, widen the investigation, and change who has to approve the next step. If an agent is making that classification and routing the case at the same time, the effect is immediate.
Take a simplified batch scenario. A deviation is logged, and the agent retrieves the approved SOP, the product specification, prior approved dispositions, and the in-process manufacturing data. Each source is controlled, yet the boundary between major and critical is not expressed in exactly the same way across them.
If the agent reconciles those sources and chooses major, the case can continue under the lighter path and the quarantine hold tied to critical never starts. Software may be operating as designed, with an intact record, valid signature, and complete audit trail. What remains unresolved is whether the agent acted on the definition the quality organization actually authorized.
Where validation stops
Traditional GxP assurance is strong because different controls protect different parts of the process. Computer system validation establishes intended software behavior. Part 11 protects records and signatures. ALCOA+ protects data integrity, while quality systems govern procedures, departures, and corrective action. That foundation does not change with agentic AI.
Before agentic workflows, a qualified professional typically absorbed inconsistencies among systems and documents before action followed. An agent can now perform part of that reconciliation inside the workflow itself. A laboratory system may hold one part of the case, an electronic batch record another, and the quality system the governing procedure; all can be valid without drawing a judgment boundary in exactly the same place.
My research calls the working meaning carried into that decision the Operational Interpretation. The control question arrives just before that meaning changes the path of the case. In the reference architecture, the Reasoning Layer is the observable runtime surface where the deployment can expose enough evidence to evaluate what the agent is carrying forward. The approach does not depend on seeing inside the model; it works with orchestration records the system already makes available to governance.

Figure 1. Six Clean Signals, One Invisible Risk. Source: Doyle-Spare (2026).
Figure 1 makes the gap concrete. Validation can confirm a qualified application, Part 11 can support the record and signature, ALCOA+ can support the data, and the quality system can show that the required steps ran. None of that establishes which definition of critical the agent actually carried into the routing decision.
Why this is not model drift
Model drift asks whether model behavior or predictive performance changes over time. This failure can occur while the models and source systems remain healthy. A workflow may still carry a version of a regulated term that has moved away from the organization's authorized definition.
A single decision can depart from the organization's authorized meaning without yet amounting to a broader operating pattern. The reference architecture calls that single-decision condition Runtime Semantic Divergence: the Operational Interpretation has moved away from the Reasoning Baseline for that term and decision class. When the same kind of divergence repeats without being contained, it becomes Agentic Workflow Drift.
Similar cases appear outside manufacturing. In pharmacovigilance, an autonomous workflow may combine a case narrative, safety-database conventions, the current product label, and MedDRA terminology before deciding seriousness or expectedness. Clinical workflows may reconcile a protocol, amendments, and site practice before making an eligibility determination. Batch disposition, CAPA priority, complaint routing, and out-of-specification investigations raise comparable control issues.
The tolerance also has to fit the decision. Deviation criticality is not the same judgment as subject eligibility or pharmacovigilance seriousness, so each decision class needs its own calibrated boundary.
The regulatory picture
Regulatory work is becoming more explicit about AI credibility, context of use, human oversight, and risk-based assurance. Those developments strengthen the surrounding control environment. They do not yet make the agent's per-decision interpretation a named governance object.
FDA's January 2025 draft guidance for drug and biological products describes a "risk-based credibility assessment framework" tied to a model's context of use. FDA and EMA's January 2026 principles address good AI practice across drug development. Europe's 2025 consultation package brought artificial intelligence directly into the GMP discussion through draft Annex 22 alongside the proposed Annex 11 revision, while ICH E6(R3) reinforces risk-proportionate quality management in clinical trials.
For an agentic workflow, timing is the part that still matters most. Before the case moves downstream, the organization needs evidence of which interpretation is driving the action and whether that interpretation is authorized. That is the moment to check the decision, while the workflow can still be stopped or routed for review.
Putting the control at the decision point
In practice, the organization first needs a governed definition it is willing to stand behind. In the architecture, that is the Reasoning Baseline for the decision class. When the agent is ready to act, the workflow exposes the interpretation it has assembled and the Semantic Deviation Index (SDI) compares the two. A separate Deterministic Gate then decides whether execution proceeds, proceeds with a recorded flag, is held for qualified review, or is placed on mandatory hold.
Those functions sit inside the Semantic Control Plane (SCP), but they do different jobs. SDI measures the distance from the approved Baseline; the Gate makes the authorization decision. A per-decision record then preserves the Baseline in force, the interpretation presented for execution, the measured divergence, and the Gate outcome so the decision can be reconstructed later.

Figure 2. Runtime Governance Enforcement at the Reasoning Layer. Source: Doyle-Spare (2026).
That sequence only works if the Gate acts before execution authority is emitted, while a divergent case can still be stopped. Post-execution review remains important for investigation and assurance, but by then the action may already have moved downstream.
That timing also keeps human review in the right place. A decision that remains within the calibrated tolerance for its class can continue, while one that approaches or crosses the threshold can be routed to a qualified reviewer before execution. Accountability stays with the organization without turning a person into the approval point for every routine agentic decision.
Validation is still the foundation
Seen in that light, the proposal adds to GxP validation; it does not displace it. Regulated organizations still need trustworthy software, controlled procedures, intact records, reliable data, qualified oversight, and defensible audit evidence. Agentic AI adds one more requirement at the point where those controls are turned into action: the organization must know which authorized definition governed the decision.
The unresolved handoff is the reasoning-layer control gap. Closing it means comparing the Operational Interpretation with the authorized Baseline before execution, while the organization still has the option to stop the case or send it for review.
A workflow can be technically valid and still carry the wrong institutional meaning. Once an autonomous system can turn interpretation directly into regulated action, that interpretation has to be governed with the same seriousness as the systems, records, and procedures around it.
How to Cite
Canonical architecture: Doyle-Spare, Maureen. Agentic AI Systems Governance: A Runtime Reference Architecture for the Reasoning Layer and the Semantic Control Plane in Regulated Financial Institutions. https://doi.org/10.5281/zenodo.20749051
Doyle-Spare, Maureen, Agentic Workflow Drift in Life Sciences: Extending the Reasoning-Layer Risk Taxonomy to GxP-Regulated Pharmaceutical and Biotechnology Operations. Available at SSRN: https://ssrn.com/abstract=7301980
About the Author
Maureen Doyle-Spare
Maureen Doyle-Spare is an enterprise governance architect and independent researcher in AI governance, specializing in agentic AI oversight and cybersecurity for autonomous and multi-agent systems. She is the originator of an agentic AI governance taxonomy built on a set of named constructs, developed for the layer above conventional AI controls, where autonomous agents interpret business meaning across fragmented enterprise systems and commit to it before acting.
Her research develops a runtime governance architecture and a foundational risk and threat taxonomy for the reasoning layer, the pre-execution stage at which an agent’s Operational Interpretation is formed and evaluated. The named constructs include Agentic Workflow Drift and Subversion (AWD/AWS), the Semantic Layer Integrity Attack (SLIA) as a cyber threat class targeting the reasoning-layer attack surface, the Semantic Control Plane (SCP) as a runtime governance mechanism, and the Semantic Deviation Index (SDI) as a measurement standard for semantic divergence. Her central thesis is that agentic systems do not fail the way traditional models fail: an agent can execute flawless steps against a meaning no institution authorized.
Her work spans behavioral visibility, pre-execution oversight, adversarial resilience, autonomous systems evaluation, and safe deployment, and policy crosswalks against the NIST AI RMF, MITRE ATLAS, STRIDE, ISO/IEC 42001, and the EU AI Act. She has provided public comments across multiple NIST AI initiatives, including the AI RMF, CAISI, COSAiS, NCCoE, and TEVV. Her working papers are available on SSRN and Zenodo.
Capstone reference architecture: https://doi.org/10.5281/zenodo.20749051
ORCID: https://orcid.org/0009-0009-6655-1394
SSRN Author Page: https://papers.ssrn.com/Sol3/Cf_Dev/AbsByAuth.cfm?per_id=10836296
ResearchGate: https://www.researchgate.net/profile/Maureen-Doyle-Spare/research
LinkedIn: https://www.linkedin.com/in/maureendoylespare/
GitHub: https://github.com/maureendoylespare/maureendoylespare
INTELLECTUAL PROPERTY: © 2026 Maureen Doyle-Spare. All rights reserved. No part of this article, including its text, figures, tables, and images, may be reproduced or distributed without the author’s prior written permission.
Maureen Doyle-Spare is an enterprise governance architect and independent researcher in AI governance, specializing in agentic AI oversight and cybersecurity for autonomous and multi-agent systems. She is the originator of an agentic AI governance taxonomy built on a set of named constructs, developed for the layer above conventional AI controls, where autonomous agents interpret business meaning across fragmented enterprise systems and commit to it before acting. Her research develops a runtime governance architecture and a foundational risk and threat taxonomy for the reasoning layer, the pre-execution stage at which an agent’s Operational Interpretation is formed and evaluated. The named constructs include Agentic Workflow Drift and Subversion (AWD/AWS), the Semantic Layer Integrity Attack (SLIA) as a cyber threat class targeting the reasoning-layer attack surface, the Semantic Control Plane (SCP) as a runtime governance mechanism, and the Semantic Deviation Index (SDI) as a measurement standard for semantic divergence. Her central thesis is that agentic systems do not fail the way traditional models fail: an agent can execute flawless steps against a meaning no institution authorized. Her work spans behavioral visibility, pre-execution oversight, adversarial resilience, autonomous systems evaluation, and safe deployment, and policy crosswalks against the NIST AI RMF, MITRE ATLAS, STRIDE, ISO/IEC 42001, and the EU AI Act. She has provided public comments across multiple NIST AI initiatives, including the AI RMF, CAISI, COSAiS, NCCoE, and TEVV. Her working papers are available on SSRN and Zenodo. Capstone reference architecture: https://doi.org/10.5281/zenodo.20749051 The Agentic AI Governence Playbook: https://www.maureendoylespare.com/ ORCID: https://orcid.org/0009-0009-6655-1394 SSRN Author Page: https://papers.ssrn.com/Sol3/Cf_Dev/AbsByAuth.cfm?per_id=10836296 ResearchGate: https://www.researchgate.net/profile/Maureen-Doyle-Spare/resear