Trust-Testing AI: Why Responsible Implementation is Not Enough

by Jenil Shah

The year is 2050, and no one talks about artificial intelligence anymore. The term is as dated as dial-up internet. The technology is embedded, dependable and a largely invisible part of the infrastructure of healthcare. A doctor arrives for her shift in an emergency department, and before she sees her first patient, an intelligent clinical system has reviewed the incoming cases, identified those at greatest risk, prepared the relevant histories and suggested the next course of action.

She reviews a prescription recommended by the system, and just before approving it, she briefly pauses. What if it is wrong? What if it has missed something? Then she accepts the recommendation and moves to the next patient. She’s right to hesitate, and that moment of doubt is evidence that the relationship between human judgement and technology is working exactly as it should.

Because at what point is it safe to completely trust an artificially intelligent system?

Trust is infrastructure

Our lives depend on trust. We allow strangers to fly the aircraft we board, perform the surgery we need, manage our savings and make decisions that affect our rights and livelihoods.

But we do not trust these systems simply because they exist. Confidence is built through standards, scrutiny, professional accountability, regulation and repeated evidence that the system works. Trust is not a feeling we apply to something at the beginning; it is a judgement we continue to make over time. That distinction matters as AI becomes more deeply embedded in how organisations operate.

Writing changed how we remembered the past, the internet changed how we found information and social media changed how quickly we could absorb and repeat other people’s opinions. AI is beginning to change how we reason, decide and exercise judgement.

It is increasingly becoming the first place people turn when they need an answer. In many organisations, it is already influencing which risks are prioritised, which patients are escalated, which applications are approved and which options decision-makers see. That should not automatically frighten us, but it should give us pause to understand how the inputs affect the outputs if every single decision isn’t scrutinised. Our responsibility is not only to build increasingly capable systems; we have to build systems that remain worthy of the authority we give them.

Responsible implementation is not enough

Most organisations still assess AI at particular points in time, such as procurement, implementation and deployment. Then the system goes live, and issues can arise because deployment is often treated as the end of the assurance process when it is actually the beginning of organisational responsibility.

Traditional software was comparatively predictable, its behaviour generally changed when somebody altered the code. AI-enabled systems are different, and their performance can be affected by changing data, new users, software updates, altered workflows, new policies and shifts in the environment in which they operate.

Once users become more confident in the system, they either learn when to challenge it or, on the flip side, they may gradually stop questioning it at all. A control that was effective when the system launched may be inadequate two years later. Models that performed well against one population may become less reliable as the data changes. A recommendation that was originally treated as one source of information may slowly become the decision itself.

That is why organisations need to move beyond asking, ‘Was this AI implemented responsibly?’ They also need to ask, ‘Does it continue to deserve our trust?’ This is the purpose of trust-testing AI: to get honest answers to these questions.

What is trust-testing AI?

At Tektology we believe trust-testing AI is the continuous discipline of assessing whether an AI-enabled system remains deserving of the trust placed in it throughout its life. It is not another compliance exercise or a final gateway before launch, it is an organisational capability.

Before development begins, it asks whether AI is the right response to the problem. During design, it challenges the assumptions being made about data, users, outcomes and risk. At the implementation stage, it examines whether the technology fits the environment in which it will operate. After deployment, it continues to test what is happening in practice:

  • Is the system behaving as intended?

  • Are people using it in the way the organisation expected?

  • Is it improving decision-making or simply accelerating existing weaknesses?

  • Are new risks emerging?

  • Are users maintaining an appropriate level of judgement, or has familiarity been mistaken for reliability?

These questions cannot be answered once. They must be revisited as long as the system continues to influence human decisions.

The three foundations of trust

Trust in AI rests on three connected foundations: governance, technical integrity and operational discipline.

Governance

Every trusted system needs clear accountability. Someone must be responsible for deciding what the system is intended to do, what it should never do and where human judgement must remain decisive. There must be clarity about who owns the risk, who has the authority to intervene and how decisions will be made when values compete.

AI governance is rarely a matter of choosing between an obviously right and obviously wrong answer. More often, organisations are balancing speed against scrutiny, efficiency against fairness, personalisation against privacy, and innovation against safety.

Those choices cannot be delegated to the technology.

Technical integrity

A system cannot be trustworthy if the data is poor, the model is unreliable or the technology is insecure. Organisations need to know whether the system performs as expected, where its limitations lie and how its outputs change over time. But technical testing cannot remain confined to a controlled environment.

The real test is what happens when AI technology meets the complexity of an organisation: imperfect data, stretched teams, inconsistent processes, competing priorities and real people making decisions under pressure. Applying sophisticated technology to a broken process does not necessarily fix the process. It may simply allow the organisation to make the same mistakes faster and at greater scale.

Operational discipline

Even a well-governed and technically sound system can fail in practice. People may misunderstand what it can do, use it inconsistently, ignore its recommendations when they should pay attention, or rely on it when they should challenge it.

Operational trust depends on training, feedback, human oversight, incident review and the willingness to learn from what happens after deployment. It also requires organisations to pay attention to behaviour:

  • Are people still exercising judgement?

  • Do they understand why the system has reached a conclusion?

  • Can they identify when something does not look right?

  • Do they know how to raise a concern, and do they believe that concern will be taken seriously?

Without that discipline, human oversight can quickly become a phrase in a policy rather than something that happens in reality.

Trust behaves like a relationship

Organisations often speak about trust as though it were a status that can be achieved. In my opinion, it is more useful to think of it as a relationship.

Relationships strengthen through consistency, transparency and care. They weaken through neglect, and the same is true here. When models drift or data changes, and people join and leave, meaning organisational priorities move, the organisation as a whole can stop noticing the assumptions built into the systems.

Controls become routine, which turns into complacency, and that becomes risk. Trust in systems rarely disappears in one dramatic moment; it decays quietly while everyone continues to behave as though it remains intact. Trust-testing is designed to interrupt that process.

It creates the habit of asking whether the confidence placed in a system is still supported by evidence.

The lifecycle belongs to the technology. Responsibility belongs to the organisation.

Trust does not follow the AI project plan timelines. It must be earned, challenged and renewed for as long as the technology continues to shape decisions. This does not mean every recommendation must be manually recreated by a human. That would defeat much of the purpose of using the technology!

The ambition should be to create systems that allow people to focus their judgement where it adds the most value.

Returning to the imaginary future in 2050 and the doctor. Perhaps she should be able to approve the prescription and move on without checking every calculation behind it. But that confidence should not come from blind faith or simple familiarity. It should exist because, over decades, people have tested the assumptions, examined the outcomes, challenged the failures, improved the controls and maintained clear accountability for the system.

Behind her single decision to approve the prescription should sit an entire infrastructure of trust. The future we should strive for is not one in which humans stop thinking. It is one in which technology has been designed, governed and tested well enough to support better human thought.

Trust-testing AI is not about slowing innovation. It is about recognising that once we believe a system is trustworthy, patients still have to be protected, so our trust has to be repeatedly earned.


Jenil Shah is a Senior Consultant at Tektology, based in Melbourne. He was previously at EY-Parthenon, working on strategy design and execution across healthcare, retail, agri-business, telecoms and manufacturing; including a global carve-out separation in the healthcare sector spanning 25 APAC markets. Before consulting, he spent time as a Project Engineer at Adani Power. His work at Tektology focuses on the operational and governance disciplines that turn ambitious digital and AI programmes into ones that actually deserve the trust placed in them.

Next
Next

Three Problems, One Discipline: The Case for Enterprise Transformation Assurance