Skip to main content

Australian defence · Decision support

Evaluating AI decision support for Australian defence

AI decision support helps people assess information, compare options and explain the basis for a decision. For Australian defence teams, evaluating that capability means looking beyond a convincing answer: examine the evidence it uses, the context it retains, the authority it respects and the record another person can review.

  • Define decision advantage through a specific user task and a measurable evaluation.
  • Assess human authority, changing context and uncertain evidence alongside answer quality.
  • Keep the results of a synthetic evaluation tied to the conditions actually tested.

Define the decision advantage you want to evaluate

Australian Defence uses decision advantage to describe making better decisions faster. Its January 2026 announcement of ASCA research contracts identifies machine reasoning, automated data integration and artificial intelligence as relevant areas. That gives teams a useful starting vocabulary; a project still needs a precise account of the decision it supports.

Start with a named user and an existing task: preparing a logistics briefing, reviewing an engineering evidence pack or reconstructing the basis for an earlier assessment. Record the information available, the decision deadline and the consequences of an incomplete or incorrect recommendation. This turns an abstract AI demonstration into a question that an evaluator can answer.

Choose measures before selecting impressive examples. Consider source accuracy, missing-information detection, reviewer effort and the time needed to reach a supported conclusion. Faster output is useful only when the team can establish what the output supports and what remains unresolved.

Defence: ASCA research supporting Decision Advantage

Make human authority explicit

Defence’s published Policy Settings for Responsible Use of Artificial Intelligence in Defence describes individual accountability, proportionate controls and assurance throughout the technology lifecycle. Treat those as requirements to examine in the workflow, rather than qualities established by an AI product label.

Write down which outputs are drafts, who may accept them and where approval is recorded. Separate permission to read information from permission to make a recommendation or pass an accepted result to another system. A fluent explanation should not expand any of those permissions.

Human-machine teaming needs an interface people can use under the conditions being evaluated. Ask whether a reviewer can inspect the source, recognise uncertainty, request more evidence and correct an earlier decision. Name the person responsible for resolving an unclear authority boundary before the evaluation begins.

Defence: Policy Settings for Responsible Use of AI

Ask for evidence at each decision boundary

The following questions are a practical starting point for a supplier discussion or a bounded evaluation. Agree what would count as a satisfactory result and retain examples of failures as well as successes. The table is an evaluation aid, not an official Defence checklist.

Questions for evaluating an AI decision-support workflow
QuestionEvidence to request
What informed this recommendation?Source identifiers, relevant passages or observations, versions and unresolved conflicts.
What changed since the earlier assessment?A record of changed inputs and whether they caused a previous conclusion to be reviewed.
Who can accept or challenge the result?Visible review responsibilities and a record of approval, correction or escalation.
What happens when information is insufficient?An example that identifies the missing basis and leaves the decision open for review.
Does the record survive a handover?A second reviewer reconstructs the decision from the retained evidence and history.
What was actually tested?The evaluated configuration, comparison method, cases, results and known limits.

Test persistent context and changing evidence

A useful first answer does not establish continuity across a long-running review. Earlier constraints may be omitted from a later prompt, a source may be superseded, or a handover summary may preserve a conclusion while losing its conditions. Include those changes in the evaluation instead of testing only isolated questions.

Check whether the system distinguishes an approved source update from an unsupported statement that a rule has changed. Ask a new reviewer to continue the task, then inspect whether the earlier caveats and unresolved questions are still visible. Persistent memory is valuable when its contents remain attributable, reviewable and correctable.

ToM places model reasoning within a persistent, governed architecture. Its public framework description explains the role of context, evidence, operating constraints and decision records. An evaluation should test those behaviours in the intended workflow rather than infer them from the model’s context-window size.

Explore the ToM framework

Design a balanced test and evaluation exercise

Use a controlled, appropriately shareable dataset and an agreed reference assessment. Include ordinary cases, incomplete evidence, conflicting revisions and changes of reviewer. Give the assisted and comparison workflows equivalent information so that the result measures the proposed assistance rather than a difference in access.

Record mistakes in both directions: an unsupported recommendation that proceeds and a well-supported task that is unnecessarily held. Measure the work needed to investigate either outcome. A system that stops every task can appear cautious while providing little practical decision support.

Defence’s account of its multinational AI Strategic Challenge identifies trust, responsible AI and human-machine teams as research concerns. That is relevant context for assessing the combined performance of the user, interface and AI, rather than judging the generated answer alone.

Defence: AI technology rises to the challenge

Inspect the decision record in ToM Defense Resolution

ToM Defense Resolution provides an operator-supervised decision workspace. Its public walkthrough shows scenario context, the basis for a proposed response, the decision reached and the recorded outcome. These surfaces let an evaluator ask how evidence, uncertainty and declared authority remain visible through a decision cycle.

The evaluation shown on the product page uses simulated observations and a synthetic environment. Resolution assesses permitted options, including no action, and can select, hold or block within that environment. Operational command-and-control and field execution are separate integration boundaries. The useful evidence is the behaviour and record shown in that defined setting.

View the ToM Defense Resolution product walkthrough

Read research results with their test conditions

Newport Resonance’s Governed Embodied Action study examines a separate supervisor between an LLM planner and a simulated robot. It compares the same sampled plans with and without that supervisor across two planner classes, using separate code to evaluate physical outcomes.

The paper reports fewer rule-violating events under its governed condition and also records unnecessary interventions and a language-only failure case. The evaluation was conducted by the authors in one simulator with small samples. It supports a bounded examination of supervision; it does not establish certification or performance in a different defence environment.

For a new evaluation, ask which parts of that evidence transfer to the proposed task and which need fresh testing. Keep the product demonstration, the research result and the next evaluation claim distinct so a reviewer can follow the basis for each.

Read Governed Embodied Action and its evidence boundaries

Assess Australian capability and long-term control

Defence’s Sovereign Defence Industrial Priorities include test and evaluation, certification and systems assurance. For an Australian AI supplier, that language points to a concrete discussion about the evidence, skills and support needed throughout a capability’s life.

Ask where development and support occur, who controls the relevant intellectual property, and which external models or services the proposed configuration depends on. Establish who can inspect the decision record, maintain the system and authorise changes. Document the data-handling arrangement for that configuration rather than assuming an Australian business address determines every dependency.

A useful next step is a scoped evaluation brief: the user task, permitted information, decision responsibilities, comparison cases, acceptance questions and deliverables. It provides a common basis for an engineering discussion with a supplier, research partner or evaluator.

Defence: Sovereign Defence Industrial Priorities

Prepare a defence AI evaluation

  • The user task and intended decision advantage are specific enough to evaluate.
  • Information access, review responsibility and approval authority are recorded.
  • Cases include changed context, incomplete evidence and a reviewer handover.
  • The comparison uses equivalent information and records both missed issues and unnecessary holds.
  • Results identify the tested configuration, known limits and questions requiring further evidence.
  • Data handling, support responsibilities and external dependencies are clear for the proposed evaluation.

ToM Defense Resolution

Explore the operator-supervised decision workspace and discuss a bounded evaluation of evidence, authority and reviewable outcomes.