Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Home
Resources

How to Judge a Delivery Partner's Use of AI: What Government Finance Teams Should Actually Expect

May 11, 2026
• 5 min read

If a delivery partner tells you they use AI, don't ask how much AI they use.

Ask what has actually changed.

Has delivery become faster? Is quality more consistent? Can they do more within the same budget? Have they reduced repetitive manual work? And just as importantly: can they explain what information goes into their AI tools, how the output is checked and who remains accountable for the result?

For a government buyer, both sides of that equation matter.

AI use should produce a benefit. It also needs to be controlled.

Why is this different for government?

Government buyers have always had a higher bar for accountability.

Value for money remains the core rule of Commonwealth procurement, but price is only part of it. Quality, fitness for purpose, supplier capability, risk and whole-of-life cost all matter.

The same principle applies to AI.

A supplier saying that AI has made them more efficient is not, by itself, evidence of better value. You need to understand what changed and whether that change benefits the engagement.

There is now another consideration as well.

Australian Government guidance specifically addresses procurements where a supplier uses AI in providing services. The Digital Transformation Agency's guidance and AI model clauses deal with matters such as approval of AI use, treatment of government information, quality assurance, record keeping, security and accountability.

This is not limited to buying an AI system. It can apply where a consultant uses AI as part of delivering otherwise conventional services.

So the question is no longer simply does this supplier use AI?

It is do they use it well?

More AI isn't necessarily better

There are two weak positions.

The first is a partner that has made no serious assessment of where AI could improve the way it works.

The second is a partner that wants to put AI into everything.

Neither tells you much about whether the engagement will be better.

There are plenty of sensible uses already. Depending on the work, AI can assist with code development, documentation, data analysis, test preparation, requirements analysis, reconciliation, research, first drafts and finding patterns in large amounts of information.

But suitability depends on the task, the information involved and the consequence of getting it wrong.

For some work, AI can remove hours of repetitive effort. For other work, the effort required to validate the result can outweigh the saving. And there will be situations where contractual, security or information-handling requirements mean a particular AI tool should not be used at all.

A mature partner should be comfortable with all three answers:

we use it here;

we use it here, but with these controls;

and

we don't use it there, and here's why.

How do I tell whether the AI story is real?

The tell is specificity.

A partner leading with the technology will usually talk about what generative AI, copilots or agents can do in general.

A partner leading with the work will tell you what has actually changed in its delivery process.

For example:

  • a first draft of technical documentation is now produced from approved project material and then reviewed by the consultant responsible for it
  • repetitive test cases are generated more quickly, while test results and exceptions remain subject to human review
  • code can be developed and checked faster, but still goes through the team's normal source control, review and testing process
  • large volumes of project material can be searched or compared more quickly without outsourcing the final judgement to the model

Those are claims you can examine.

"We are an AI-enabled consultancy" tells you very little.

What questions should I actually ask?

A government buyer does not need to conduct an interrogation of every tool a supplier owns.

A handful of direct questions will tell you a lot.

"Where are you actually using AI in our engagement?"

You are asking about the delivery method, not the supplier's general AI capability.

A weak answer describes what AI can do.

A good answer names the parts of your engagement where it is used, what it assists with and what remains unchanged.

It should also be perfectly acceptable for the answer to be "we are not proposing to use AI for that part of the work."

"What information are you putting into those tools?"

For government, this may be the most important question.

You should know whether project information is being processed through AI tools and, if so, under what conditions.

That does not mean every engagement needs a lengthy AI security assessment. The response should be proportionate to the risk.

But a supplier should be able to explain which tools it permits staff to use, whether those tools are enterprise-managed, what client information can and cannot be entered, and how its approach complies with the security, privacy and contractual requirements of the engagement.

"We just use ChatGPT" is not an adequate information-handling policy.

"How do you know the output is good enough?"

AI can accelerate work. It can also produce convincing mistakes.

The appropriate control depends on what the AI is doing.

A draft meeting summary and a calculation feeding a financial model do not carry the same risk. Neither do a coding suggestion and an automated decision affecting a member of the public.

A good partner should be able to explain where review is required, what is tested, how errors are detected and when a person needs to make the judgement.

The point is not that a human must manually redo everything the AI has done.

The point is that the supplier understands where human oversight matters.

"Who is accountable for the result?"

This should have a simple answer.

The partner is.

Using an AI tool does not transfer responsibility to the model vendor.

If a consultant submits a design, builds a model, writes code or produces an analysis, the normal professional responsibility for that work should remain clear.

That becomes more important as suppliers begin using AI agents capable of undertaking longer sequences of work with less direct intervention.

The technology may be more autonomous. Accountability should not be.

"What has genuinely improved?"

This is where the conversation moves from governance back to value.

Ask what is now faster, more consistent or less labour-intensive.

Has testing changed? Is documentation produced earlier? Can more alternatives be assessed during design? Are developers spending less time on repetitive code? Can the same team cover more scope?

If the answer is "not much yet", that can be more credible than an exaggerated productivity claim.

AI does not create the same benefit in every type of work.

What matters is whether the supplier understands where it is actually making a difference.

"Has that improvement changed what we receive for our money?"

This remains a fair question.

If AI has materially improved a partner's productivity, the client should see some benefit.

That does not necessarily mean a lower hourly rate.

It could mean fewer hours, a lower fixed price, more work within the same budget, a shorter delivery period, better testing, better documentation or more time spent on the difficult parts of the engagement.

The point is not to squeeze every efficiency out of the supplier.

It is to make sure the AI story and the commercial story make sense together.

If a partner claims a substantial productivity improvement but the same work still takes the same team the same amount of time at the same cost, it is reasonable to ask where the benefit went.

What should a good answer sound like?

Concrete, measured and internally consistent.

It names the tools or types of tools being used. It explains which tasks have changed. It identifies the information-handling boundaries. It tells you where review is required and who is accountable.

And it should connect the use of AI to an outcome you care about.

A credible answer does not need to claim that every part of delivery has been transformed.

In fact, caution in the right places is a positive sign.

Current Australian Government guidance takes a risk-based approach to AI. It asks agencies to consider issues including reliability, safety, privacy, security, transparency, human oversight and accountability. The same thinking provides a useful way to assess suppliers.

The mature position is not "AI everywhere".

Nor is it "AI nowhere".

It is knowing where AI makes the work better, knowing where it doesn't, and having appropriate controls around the difference.

Where does this leave a government buyer?

You do not need to become an AI specialist to have this conversation.

You do need enough visibility to understand how your supplier is working on your behalf.

Ask where AI is being used. Ask what happens to your information. Ask how the work is checked. Ask who owns the result. Then ask what has actually improved.

A good delivery partner should be able to answer those questions without retreating into a technology presentation.

If they can, you have learnt considerably more than whether they "use AI".

You have learnt something about how they manage delivery, risk and accountability — which is ultimately what you are buying.

If the content doesn’t load, open it in a new tab .

Related Resources