After a year of using artificial intelligence in QA/QC, I discovered that the real test is not for the machine
Listen to this article
Read by Anchor
About a year has passed since I began actively using artificial intelligence in quality assurance and quality control (QA/QC) work.
This was not simply a matter of asking it to polish an email or summarize a text.
I mean deploying it alongside inspection reports, technical specifications, procedures, field notes, non-conformance reports, audits, inspectors working across multiple sites, a client awaiting an answer, and a technical decision that leaves no room for guesswork.
At first, I was impressed.
A task that used to take half an hour was finished in minutes.
Hurried notes from an inspector could be organized.
A poorly phrased report could be improved.
A lengthy document could be summarized.
A technical email could be refined before going to the client.
I asked myself then:
What were we spending all that time on?
Over time, however, the question shifted.
It became:
Could saving time lead us to sacrifice quality itself?
After a year, I reached a conclusion that may matter more than anything I learned about prompting or choosing tools:
Artificial intelligence does not test the technology alone. It tests the quality culture under which we operate.
If your system is disciplined, you can increase its efficiency.
If your culture is built on letting things slide and getting the report out at any cost, artificial intelligence simply produces the problem faster and in a cleaner format.
This is precisely where the concern begins.
The first time I doubted a polished answer
One of the quickest lessons I learned was that polished phrasing does not equal accurate information.
Once, while reviewing a technical point, I used the model to help identify the relevant reference.
I was impressed by how well organized the answer was.
The logic appeared sound.
The terminology was correct.
The clause number was cited with complete confidence.
Had I been looking only for a quick answer, it would have been easy to rely on what the model produced.
Yet I felt doubtful and opened the reference.
That was where the problem appeared.
The cited clause did not support the claim in the way it was presented.
For me, that moment proved more valuable than ten lectures on hallucination.
It showed me the risk in practical terms.
The problem is not always that artificial intelligence produces an obvious error.
At times, it deliversa mistake that looks thoroughly professional.
An obvious error is easy to catch.
An error written in clean prose, backed by a confident explanation, can easily pass if the reviewer wants to believe it.
From that point on, I adopted a personal rule:
Never treat the confident tone of an artificial intelligence model as a substitute for verifying the primary source.
In quality assurance and control, the approved standard governs.
Not the look of the answer.
Not the speed of its output.
Not how convincing the model sounds.
Assistant or expert?
I return to this question constantly.
If a new engineer joined my team, it would be normal to ask them to draft a checklist, summarize requirements, or review a document and record notes.
Yet would I approve the outcome without review?
Naturally not.
Here I noticed a strange contradiction.
We often review the work of an engineer with years of experience, yet accept an answer generated by a model in ten seconds simply because it looks right.
In quality work, looking right is never a standard to rely on.
Information has eitherbeen verified, or it has not.
That is where I defined the role of artificial intelligence for myself.
It is one of the fastest assistants I have worked with.
It does not tire of rephrasing.
It does not object when asked to review a text a second or third time.
It can organize large volumes of information quickly.
Yet there is a fundamental difference between it and any accountable person in the organization:
It holds no signature to bear responsibility for the outcome.
I do.
Quality systems already know how to handle AI
When I first began reading about AI governance, many terms felt familiar:
Traceability.
Verification.
Authority.
Controls.
Human review.
Risk management.
I asked myself:
Have we not been doing this for years?
Consider how it works.
We have procedures.
We have records.
We have traceability.
We have non-conformance management.
We have corrective action.
We have audits.
We may not need to invent an entirely new quality philosophy for artificial intelligence.
We simply need to refuse to exempt it from the standards we apply to everything else.
If artificial intelligence enters a process, I must know its exact starting point.
If it helps draft technical material, it must be clear who reviewed that output.
If a mistake occurs, we cannot simply say the model erred.
We must ask:
Why did this error slip through?
Was the input unclear?
Did the user ask for an inference with no reference provided?
Was the review conducted properly?
Was there a genuine human control point, or merely an assumption that someone surely checked it?
These are not merely IT questions.
These are quality questions.
The role of QA/QC may be larger than we imagine in the age of artificial intelligence.
Prompts matter, but they are not magic
Prompt engineering certainly matters.
Yet in quality control, I do not see it as an obscure skill or a collection of magic phrases.
To me, a good prompt resembles a good inspection instruction.
You define the task.
You define the reference.
You define the scope of work.
You specify what cannot be assumed.
You request clarification on points that require human review.
If you simply write:
“Prepare a welding checklist.”
You have provided incomplete instructions.
Which revision?
Which scope?
Which welding process?
Which project specification?
Do you need a general checklist or one tied to a specific activity?
Is the model permitted to generate acceptance criteria, or only extract what appears in the provided standard?
The problem is not always a unreliable model.
Often the issue is that the user does not define the task precisely enough to produce a controllable outcome.
At that stage, prompt drafting turns from a mere operational skill intoa method of risk control.
Where real limits begin: the welding procedure specification
Take a welding procedure specification (WPS) as an example.
Can I use AI to prepare a review checklist?
Yes.
Can it summarize certain requirements or compare existing data against specific criteria?
Yes.
Can it help identify a missing field or a point that warrants attention?
Certainly.
Yet would I hand it a WPS and ask it to:
“Approve or Reject?”
Here my answer is clear.
No.
Not because the model cannot output the words approve or reject.
It can.
The issue is that generating a word is not making a decision.
A decision depends on individual competence, designated authority, the applicable code, project specifications, supporting documentation, and professional accountability.
This is not a question of technical capability.
It is a matter ofaccountability.
Artificial intelligence can assist in preparing a decision.
It cannot own the decision.
Product release raises higher stakes
The same principle applies to product release.
Artificial intelligence can summarize the status of inspection reports.
It can organize open punch items.
It can review a punch list.
It can help identify documents that appear missing.
All of that is useful.
Yet a release decision does not live within a textual summary alone.
Behind it lie material certificates.
Test reports.
Non-conformance reports.
Acceptance criteria.
Hold points.
Client requirements.
Specific signatures and authorized sign-offs.
The data supplied to the model may also be incomplete from the start.
Here lies the question we must examine:
Who bears responsibility for what the model never saw?
What if the latest report was not uploaded?
What if an open NCR was omitted from the data?
What if a unapproved revision of a specification was used?
The model can analyze what it receives.
It bears no responsibility for what it lacks.
At times, the most dangerous wrong decision is one based on an excellent analysis of an incomplete dataset.
Confidentiality
This may be one of the areas requiring the most organizational maturity and awareness.
Enthusiasm for productivity easily leads users to think:
Upload the entire specification.
Upload the report.
Upload the NCR.
Let it read everything to save us time.
That sounds convenient.
Yet before pressing upload, a more critical question arises:
Do I have the right to upload this information here in the first place?
Client specifications.
Project drawings.
Material certificates.
Test reports.
Commercial data.
Supplier details.
All of these may fall under contractual obligations or internal governance policies.
Good intentions are not enough here.
An employee may simply want to finish their work faster, while simultaneously transferring proprietary information to a unapproved tool.
The intention may be good, but a breach remains a breach.
For this reason, I treat confidentiality as an integral component of the quality process itself, rather than a technical afterthought.
After one year, do I trust artificial intelligence?
Yes.
And no.
I trust it as a tool.
I trust its capacity to accelerate drafting, organization, preparation, and initial comparisons.
I trust that it can make the workday more efficient.
Yet I do not treat it as a final authority on critical technical data.
Nor as an approving entity.
Nor as a substitute for a qualified specialist.
This is not a contradiction.
I trust a measuring instrument, yet I verify its calibration status.
I trust an inspector, yet I check credentials and qualifications.
I trust a test report, yet I confirm traceability.
Why should artificial intelligence be the sole element in a quality system that we are asked to trust without controls?
The problem is rarely AI alone
After a year, this may be my most solid conviction.
If someone uses AI to fabricate missing data, the fault does not lie with the model alone.
If an engineer copies a clause number without verifying it, the fault does not lie with the model alone.
If a technical conclusion passes simply because it is well phrased, our review process is flawed, not just the technology.
Artificial intelligence reflects the culture of the organization.
Where the culture is strong, it acts as a powerful tool within it.
Where the culture is weak, it can cause something more hazardous:
Artificial intelligence does not fix a weak quality system. In some cases, it merely allows a weak system to produce better-looking documents.
That is not digital transformation. It is an old problem dressed in a new suit.
The takeaway
After a year of practical use, my questions have changed.
I no longer ask simply:
Is the team using artificial intelligence?
I now ask:
What are they using it for?
What kind of data enters the system?
What outputs do we rely on?
Who verifies them?
Which decisions must remain under strict human control no matter how capable the tool becomes?
For me, the future of AI in QA/QC is not determined solely by how smart the models are.
It is determined by the maturity of the quality systems that receive them.
I still treat artificial intelligence as a very new team member:
Fast.
Capable.
Useful.
Worth leveraging as much as possible.
Yet if it tells me regarding a sensitive technical matter:
“I am sure.”
My response remains unchanged:
Fine, show me the reference.
And even if it provides the reference?
I will open it myself.
The conclusion after a full year is neither that we should resist artificial intelligence, nor that we should surrender our work to it.
The takeaway is simpler:
Let it help you prepare the decision, but never let it own the decision.