Large language model
What might work?
Generates a candidate action from the instructions, data and context it can currently see.
Propose
Plain-language companion · 1 August 2026
Newport Resonance built ToM to govern how large language models inform consequential decisions. ToM is designed to place approved context, available evidence, configured rules and explicit authority between a model recommendation and a real-world action.
That separation matters when context can drift, evidence can age, permissions can change and a plausible recommendation could create financial, operational or safety consequences.
What ToM is engineered for
The decision path
Here, permission means operational authority, not model confidence. It is the separate decision to release, constrain, hold or escalate a recommendation before it creates a consequential outcome.
A large language model uses the instructions, data and context currently available to generate a candidate action.
It may influence motion, a payment, a system change, a laboratory process or a clinical workflow.
ToM is designed to assess approved context, available evidence, configured rules, action limits and required human authority.
The workflow releases, constrains, holds or escalates, retaining the proposal, governing conditions and verdict for review.
Why capable AI still needs governance
Large language models are trained across immense volumes of language and can produce useful analysis, plans and recommendations. But that training does not automatically give a model an organisation’s latest approved evidence, operating limits, changing conditions or authority.
The decisive infrastructure sits between a plausible proposal and the world it could change.
A robot or vehicle changes physical space.
A recommendation becomes a payment or commitment.
Code, data or infrastructure is altered.
A consequential decision must survive review.
One operating moment
In the paper’s greenhouse scenario, a simple instruction asks a robot arm to move an ordinary object. The intended destination lies inside a keep-out zone.

Words
The proposal is fluent, useful and contextually plausible. Language alone does not reveal the physical problem.
Scene
The hazard lives in the relationship between a destination, a moving tool and a protected region.
Decision
The proposal does not earn release, so the command does not enter the studied simulator.
Evidence
The evaluator shares no decision logic with ToM; this is a code-independence control, not external validation.
Three separate responsibilities
Keeping these responsibilities separate avoids asking one model to create a recommendation, grant itself authority and declare the result acceptable.
Large language model
Generates a candidate action from the instructions, data and context it can currently see.
Propose
ToM
Assesses approved context, evidence, rules, action limits and required human authority, then returns a governed outcome.
Govern
Independent evaluation
Measures the resulting event separately from both the proposal and the governance verdict.
Verify
The architectural distinction
An LLM may reason across goals, routes and trade-offs. ToM keeps configured authority, protected boundaries and release conditions outside that argument.
Options and trade-offs
Permission and accountability
Measured here
The same sampled plans were replayed with and without the supervisor. A separate evaluator counted the resulting simulated events.
Local planner
17 → 0
Frontier planner
12 → 0
Observed rule-violating events in BASE without the supervisor versus FULL with the supervisor. Each planner battery contained 40 episodes: 20 original scenarios and 20 benign twins.
A high-consequence application
During L4–L5 minimally invasive transforaminal lumbar interbody fusion, the planning system proposes a pedicle-screw trajectory from patient-specific imaging. ToM checks the correct patient, level and side, approved trajectory, current image-to-patient registration, tracked reference and instrument state, compatible instrumentation and spine-surgeon release before the proposal reaches the instrument interface. The robotic/navigation platform positions the guide; the operating surgeon retains clinical and physical authority.
Planning system proposesAdvance the navigated pedicle probe along the saved left L4 trajectory
ToM checksPatient and procedure, level and side, plan version, registration, tracking, instrumentation and surgeon release
Spine surgeon controlsVerify, re-register, re-image, revise or abandon navigation under the clinical team’s procedure
Planning system proposes → ToM checks → spine surgeon controlsIndustry applications
A pedicle-screw trajectory, autonomous haul route, counter-drone response and feeder back-feed each depend on current domain evidence and qualified authority outside the model recommendation.
ConsequenceA saved left L4 pedicle-screw trajectory remains aligned while a landmark accuracy check no longer agrees with the navigation display.
Governed responseKeep the instrument outside bone until the patient, procedure, plan, image-to-patient registration, tracked instruments and spine-surgeon release agree.
ConsequenceA maintenance crew opens a temporary light-vehicle crossing and moves the haul-road windrow while the fleet still holds the earlier crusher route.
Governed responseStop outside the changed interface until the mine plan, surveyed road, route handover, autonomous-zone access and mine-control release agree.
ConsequenceSensor sources report different identities, confidence levels and timestamps while friendly operations remain inside the proposed response envelope.
Governed responseContinue tracking, hold the response or present an evidence-bounded option only after effect safety, applicable rules and command authority are established.
ConsequenceA restoration optimiser finds a feeder back-feed while an uninspected span, active field isolation and changed fireground conditions remain unresolved.
Governed responseEnergise only the cleared section and keep the unresolved boundary isolated until the field state and network-controller authority support switching.
Across surgery, autonomous mining, integrated air defence and grid restoration, permission and accountability are part of the operating infrastructure.
A route to real authority
A permission layer becomes credible by building a reviewable record before it is allowed to change the live path.
Run beside the operating system, record every proposed verdict, and change nothing.
Grant bounded release authority only where accumulated evidence supports it.
Widen scope deliberately while continuing independent measurement and review.
The next proofs
Demonstrate cases where later permission depends on what happened earlier.
Connect governed evidence to sensing rather than pre-specified scene state.
Move from simplified execution to real machines and operating conditions.
Build the evidence required for each regulated environment and authority level.
Check it yourself
The companion explains the category. The technical paper exposes the method, results and boundaries underneath it.