SynHy Article

Safety-Critical AI Needs A Human Decision Envelope

Safety-critical AI needs a human decision envelope that defines what the system may recommend, what people must decide, when automation disengages, and how authority is restored under failure.

Decision Support Can Change Safety Without Taking Control

An AI system does not need to steer an aircraft, operate a machine, or issue a clinical order to affect safety. A recommendation can change which option a person notices, how quickly they act, what they believe is normal, and whether they retain the skill and information needed to challenge the system.

The United States is introducing new air-traffic flow-management software intended to coordinate schedules and trajectories and reduce congestion. Public descriptions emphasize decision support and strategic planning rather than autonomous control of aircraft. That distinction should be converted into an explicit operating boundary before users begin relying on the recommendations.

Authority Drifts Through Repeated Reliance

Formal policy may say a human remains responsible while interface design and workload encourage routine acceptance. If recommendations arrive faster than staff can independently evaluate them, the practical decision can migrate to the system. People may become monitors who are asked to intervene only when the situation is least familiar and time is shortest.

Automation also changes shared mental models. Users need to know what data the system sees, which objective it optimizes, where it is uncertain, and how its failure modes differ from human error. Calling the software a teammate or brain can obscure the fact that responsibility remains with designers, operators, and accountable organizations.

Failure Cost Includes Recovery Time

Safety-critical cost cannot be reduced to the probability of an incorrect recommendation. It includes detection delay, the number of affected operations, time to reconstruct state, workload transferred to people, loss of separation or safety margin, downstream schedule disruption, and the quality of the fallback system.

A practical scenario measure is affected decisions multiplied by exposure time and consequence, adjusted by detection and successful recovery. Use ranges rather than a single expected-loss number. Rare, correlated failures deserve attention because one bad input, shared model, or unavailable service may influence many decisions at the same time.

Diagnose The Real Decision Path

Map each recommendation from source data to displayed option, human review, action, system confirmation, and later correction. Identify time available, alternative information, required expertise, and whether the operator can understand why the recommendation changed. Observe real work rather than relying only on the designed procedure.

Warning signs include automatic acceptance defaults, weak indication of stale or missing data, alerts without prioritization, no recorded reason for overrides, training based only on normal scenarios, and a fallback that staff rarely practice. If users cannot continue safely when the tool is unavailable, the system has acquired operational authority regardless of policy language.

Choose The Level Of Assistance Deliberately

The system may summarize information, rank options, predict congestion, recommend a plan, execute a reversible preparation step, or carry out an action after approval. Each level changes workload and risk. A higher automation level is not inherently better; it is justified only when evidence shows the complete human-system process performs more safely.

Keep manual or conventional tools where uncertainty, consequence, or data weakness is high. Use staged deployment, shadow mode, advisory mode, restricted geography, limited traffic conditions, and explicit approval to build evidence. Avoid mixing experimental recommendations into live work without making their status unmistakable.

Define The Human Decision Envelope

The envelope should specify permitted recommendations, prohibited actions, data prerequisites, operating conditions, uncertainty limits, human role, approval point, maximum response time, disengagement trigger, fallback tool, and restoration authority. It should name who can narrow or suspend the envelope when evidence changes.

Define what the person must be able to see and do. This includes source freshness, conflicts, alternative options, downstream effects, and a clear way to reject the recommendation without fighting the interface. Record the system's suggestion, the human decision, the reason for material divergence, and the operational outcome for later learning.

An Airspace-Flow Example

Consider an illustrative decision-support tool that recommends departure sequencing to reduce congestion. The envelope permits recommendations when weather feeds, airport capacity, route restrictions, and traffic data meet freshness thresholds. A controller or traffic manager approves the plan, and the tool cannot issue clearances or create new operating procedures.

If a critical feed is stale, demand exceeds the tested range, or recommendations conflict with a safety constraint, the tool visibly disengages and the conventional process resumes. Staff practice that transition under realistic workload. Restoration requires data recovery, a known-good state, and authorization from the named operational owner rather than an automatic restart.

Measure Joint System Performance

Track recommendation accuracy, acceptance and override patterns, stale-data events, unsafe or infeasible suggestions, time to detect failure, workload, situation awareness, fallback performance, and recovery time. Compare the human-system team with the prior process across routine, degraded, rare, and adversarial scenarios.

Watch for automation bias and skill decay. A rising acceptance rate is not proof of improvement, and frequent overrides are not necessarily user resistance. Review whether the system is wrong, the envelope is too broad, training is incomplete, or the recommendation arrives without enough context for responsible judgment.

Run A Disengagement Drill

Choose one high-consequence recommendation and write the current envelope on a page. Remove one important data feed, introduce a conflicting constraint, and test whether users recognize the condition, reject the recommendation, transition to fallback, and preserve a shared operating picture. Measure time, errors, communication, and recovery.

Repair the procedure, interface, or scope before expanding automation. Repeat the drill with new staff and peak workload. Safety evidence belongs to the complete operating system: technology, people, rules, training, communication, and fallback. A strong model cannot compensate for an undefined transfer of authority.

Sources, Method, And Limits

This article was prompted by reporting on a planned U.S. air-traffic AI tool and checked against the FAA announcement describing FMDS and SMART. The FAA says the system will support traffic-flow planning and strategically coordinate schedules and trajectories. Public material does not establish every deployed function, limit, or assurance result.

The envelope is SynHy original analysis informed by the FAA Roadmap for Artificial Intelligence Safety Assurance and the FAA's Safety Framework for Aircraft Automation, including cautions about personification, workload, and assigned responsibility. It is general operational guidance, not aviation certification or safety approval.

Does This Sound Familiar?

If this article brings to mind a slow process, repeated task, or frustrating handoff in your business, let’s talk about it. We’ll help you explore what could work better.

Let’s Talk About Your Workflow