An Operating Licence for AI
Defining what an AI system has proved it can do, and how far it should be allowed to go
We have been working with a technology company developing a smartphone application for quality control in garment manufacturing. The concept is straightforward enough. A phone positioned above a workstation photographs a garment, compares selected measurements with the customer specification and tolerance, and creates a photographic record of the inspection.
The attraction is easy to see because garment inspection involves a large amount of repetitive checking, while the underlying information is already available in the form of product specifications and approved tolerances. The garment is physically present, the measurement can be repeated and an operator standing beside the workstation can examine the product whenever a result looks unusual.
This also creates a useful environment in which to establish whether the application works. Testing can cover particular garments, specified measurements, camera positions, lighting conditions and tolerances, gradually building evidence of how reliably the system performs across the range of circumstances it is expected to encounter.
A more interesting question appears once the application begins to work well.
During our discussions, fabric shade provided a good example. A system which can measure the dimensions of a garment from a photograph may appear well placed to decide whether two garments are sufficiently close in colour. Both activities involve a camera and a garment, and to the person using the application they may feel like extensions of the same inspection process.
The underlying problem is rather different. Shade assessment requires its own training data, testing criteria, lighting controls and validation. Evidence that the system can measure a sleeve accurately provides evidence about sleeve measurement, while shade requires another body of evidence before management can have the same degree of confidence.
As AI applications become successful, this distinction becomes important because confidence tends to travel further than the evidence which created it. A system may perform one task exceptionally well and gradually acquire additional functions, while users increasingly treat its general capability as established. In reality, management knows something much narrower and more useful: it knows the conditions under which particular capabilities have actually been demonstrated.
The garment application raises another issue because the operator remains able to inspect the product, repeat a measurement and override the AI judgement. This appears to provide a strong human control, especially during the early stages when operators are still learning how the application behaves.
Over time, the nature of that control can change. An operator who has watched the system make hundreds or thousands of correct decisions is likely to develop confidence in its judgement, which is exactly what we would expect from someone working with a reliable tool. Independent review can then drift gradually towards routine confirmation because experience has taught the operator that the machine is usually right.
The override remains available and the human remains present, while the practical behaviour surrounding the control has changed. That suggests that management needs evidence about the operation of the control as well as evidence about the AI itself. Override rates, disagreements, subsequent findings and patterns of routine acceptance can show whether the human judgement built into the process continues to operate in the way management intended.
The same problem becomes harder to see when the AI is working with information because there is no garment sitting on a table and no physical measurement which can immediately be repeated.
Each morning, for example, my AI assistant can prepares a report on the previous day’s website performance. It retrieves information from approved sources, compares periods, calculate movements, identify unusual changes, research possible explanations and produce a written report .
At first sight these activities sit comfortably together because they appear in the same report, yet they involve quite different forms of competence.
Retrieving yesterday’s website traffic is largely a data retrieval exercise, while calculating the percentage movement from the previous day introduces a straightforward calculation. Explaining why traffic changed requires the system to move into interpretation, where the range of possible explanations becomes much wider.
A decline in traffic could reflect search visibility, technical changes, seasonal behaviour, competitor activity, changing customer interest or ordinary variation. The AI can investigate these possibilities and collect evidence, although the strength of that evidence can vary considerably from one explanation to another.
The difficulty for me reading the report is that fluent language tends to make these differences less visible. A strongly supported conclusion and a speculative explanation can both arrive in polished English, written with equal confidence and presented within the same document. The quality of the prose tells the reader very little about the quality of the evidence behind the conclusion.
This changes the management question. The reporting assistant may have demonstrated that it can retrieve defined information accurately, calculate movements, compare periods and identify exceptions. It may also be capable of researching possible causes and bringing together evidence for a manager to consider. The degree of confidence attached to attribution then becomes part of the operating design because explanation involves a different evidential standard from retrieval and calculation.
Further capabilities create another change again. The same AI may technically be capable of modifying the website, publishing material, contacting a customer or committing expenditure. Each additional capability changes the possible consequence of an error, even though the underlying model may remain the same.
This is the point at which the analogy with aviation becomes useful.
A pilot’s licence describes more than the ability to fly an aircraft. It establishes the aircraft, operating conditions and circumstances within which competence has been demonstrated. Additional aircraft, night flying or instrument conditions bring additional requirements because competence in one operating environment provides only part of the evidence needed for another.
The aircraft operates within defined conditions as well. Its ability to fly says little about whether a particular load, weather condition or operating circumstance falls inside the envelope for which it has been approved.
The same distinction appears throughout ordinary management. An employee may be competent to prepare a payment and have a separate level of authority for approving it. An engineer may understand exactly how to make a technical change while the organisation places that change within an approval process because the consequences extend beyond technical competence.
AI brings these familiar management ideas together in a slightly unusual way because capability can expand extremely quickly. A system may acquire access to more information, stronger reasoning, additional tools and the ability to take actions across several business systems, while the organisation still has to decide how much of that capability it wishes to use and under what circumstances.
Recent incidents involving AI agents have made this issue more visible because systems sometimes encounter routes towards an objective that their designers failed to anticipate. Complex systems will continue to encounter combinations of circumstances which were absent from testing, particularly as their environments, tools and objectives become broader.
The practical management problem therefore extends beyond predicting behaviour. It includes deciding how far unexpected behaviour can travel before the system reaches a boundary established by the organisation.
The garment application provides a relatively simple example because an uncertain measurement can return to the operator standing beside the product. The reporting assistant can reach a point where the available evidence supports several possible explanations and present that uncertainty to the manager. A system with access to the website can prepare a proposed change while publication remains subject to separate authority.
In each case the technical capability of the AI can be greater than the authority granted to it. The organisation decides how much consequence can follow from the system’s own judgement.
This suggests the need for something more explicit than a general approval to use an AI application. We have started to think of it as an Operating Licence.
The licence would describe the purpose for which the application has been approved, the capabilities for which there is evidence of reliable performance and the operating conditions within which that evidence remains valid. It would also define the data and systems the application can access, the actions it can take within its own authority and the point at which another person or another approval becomes necessary.
The consequence of independent action would form part of the same definition. A system producing an internal management report operates in a different environment from one able to alter a customer website, release a payment or communicate externally, even where both systems use similar underlying AI technology.
The licence would also establish what happens when the system reaches the edge of its demonstrated competence. Some circumstances may lead to escalation, some to a request for further evidence and others to a predetermined safe response. The appropriate answer depends upon the application and the consequence attached to the decision.
Continuing evidence would matter because the licence describes an operating condition rather than a permanent judgement about the technology. Performance can be monitored, human overrides can be reviewed and changes in the model, data, task or working environment can trigger further testing.
This allows the licence to expand as evidence accumulates. The garment application might begin with particular products and measurements before additional garment types, inspection tasks or operating conditions are added. The reporting assistant might begin with retrieval, analysis and reporting before management decides that experience supports a wider range of permissions.
The reverse process is equally useful. A change in operating conditions, evidence of deteriorating performance or weakening human oversight can lead to a narrower operating boundary while the cause is understood and the evidence rebuilt.
Seen in this way, the Operating Licence becomes part of management rather than a technical document produced once during implementation. It records what the organisation has learned about the system, defines the authority that management is prepared to grant and provides a basis for deciding when that authority can expand.
AI capability is likely to continue increasing rapidly, and many systems will soon possess technical abilities which organisations have little reason to use in every application. The important decision is therefore how competence, authority and consequence should remain connected as those capabilities develop.
Aviation reached much the same position long ago. Pilots and aircraft operate within defined conditions, experience can extend those conditions and additional qualification allows the operating envelope to grow. The discipline lies in knowing where the boundary currently sits.
An Operating Licence for AI would apply the same principle to systems whose capabilities may continue expanding long after they enter service. Management can allow that capability to develop while extending authority at the pace justified by evidence of how the system actually performs.