# Is AI Only Bad at Your Job?

> We catch AI's mistakes in our own work, then trust it in unfamiliar fields. Are the answers better, or are we less able to judge them?

A developer reviewing AI-generated code can tell the difference between an application that runs and a system that is well designed. The code may do its job. The developer still spots a business rule in the wrong place, a failure case nobody considered, or a design choice that will make the next change harder.

An SEO specialist makes a similar distinction. The recommendations look reasonable, but they miss the search intent. They list things that could be done without explaining which ones matter most for that site. A dietitian, meanwhile, knows the difference between a neatly organised meal plan and one suited to a particular person.

In our own fields, we tend to give AI a qualified verdict: useful, but you need to know what to look for before you can trust the output.

Now imagine the developer getting a meal plan from AI, the dietitian relying on AI to make an application's technical decisions, and the SEO specialist using an AI-generated financial analysis to guide an investment. That caution can be harder to maintain in an unfamiliar field.

So AI is good at everyone else's job, just not yours? The question is aimed at our own inconsistency, not at those professions. Failing to see the flaws outside our field does not mean the flaws aren't there.

## You can miss an error without knowing it

When we read code, we look beyond what has been written. We may notice a missing check, wonder what happens if the same request arrives twice, or see how an easy choice today could complicate future work.

Knowing the field helps us assess an answer and notice which questions it leaves unanswered.

In an unfamiliar subject, our checks may be more superficial. Is the text clear? Is it internally consistent? Does it give detailed reasons? Does it answer the question we asked? We can judge those things. We may not know whether an unmentioned condition would change the answer.

In a financial analysis, checking that the numbers add up and checking that the starting assumptions are reasonable are separate tasks. We can repeat the calculation. Recognising why an assumption is weak may require knowledge we don't have.

“AI is much better at finance than at software” could reflect a genuine difference in performance. It could also mean we recognise the software errors and miss the financial ones. Until we have ruled out that second possibility, we should be cautious about the comparison.

AI need not perform equally well in every field. The scope of the task, the information supplied, and the ways of testing the result all vary. Expertise does not make us infallible reviewers either. Still, finding errors in one field and none in another does not, by itself, measure the difference in quality.

## Getting an answer doesn't give us the means to judge it

AI can put an analysis draft within reach when we could not have prepared one ourselves before. We can also investigate an unfamiliar subject more quickly. That is a useful capability. Receiving the output, though, does not also give us the knowledge to assess its quality.

We might understand each SEO recommendation on a list. Deciding which one would help a business acquire customers requires us to consider the site, its competitors, and the business's priorities together. A long list of possible actions leaves that decision unresolved.

In my [essay on the cost of verifying AI-generated code](/ai-made-code-cheap-verification-is-still-expensive), I described how much thinking an apparently finished application can leave to its reviewer. The problem here starts earlier: we may not know enough about the field even to recognise that a review is needed.

As it gets easier to produce work, deciding what is worth using matters more. Being able to produce more reports in the same time can help. If we cannot tell which reports can support a sound decision, our decision quality will not improve at the same pace. We will also have more material to review.

## What exactly are we checking?

“Always check AI's output” leaves a question unanswered: what should we check?

Rereading an answer in an unfamiliar field may help us understand it better. It does not guarantee that we will spot a faulty assumption. Asking AI to criticise its own answer can expose omissions too. We still have to judge whether the second response is a good critique.

Verification is useful when we have something concrete to compare the answer with. We can compare figures with actual records, open a source and read the condition attached to a claim, or observe software behaviour in a bounded test. Where our knowledge falls short, someone qualified can assess the result. If they reject an assumption, we can ask which one and why.

AI can help with that learning. We can ask it to explain terms, compare alternatives, or identify conditions that might change the answer. When we then test those explanations against sources or actual results, we gain more than an answer. Over time, we can get better at judging the work ourselves.

The same question applies when a company adds human approval to an AI workflow. How will the person approving the output recognise an error in that field? In my [article on trusting AI judgments](/why-we-trust-ai-judgments), I distinguished accepting advice from making a sound decision. A manager reading a report does not, on its own, show that its technical or financial assumptions have been tested.

## The consequences determine how much checking we need

Asking for a different dinner recipe and making a treatment decision for a health condition call for very different levels of scrutiny. We can easily move on from a meal we don't like. With treatment, discovering a mistake afterwards may not be enough.

Work presents similar differences. Reviewing an AI-generated database design in a test environment is one decision; putting it straight into a system customers use is another. Learning the terms in a legal document differs from signing an AI-drafted contract and taking on its obligations.

What we intend to do with the output should determine the checks, more than the name of the profession involved. Who would be affected if it were wrong? Could we reverse the decision? Could we independently tell whether the result was good, or would we learn only when harm occurred?

A small trial may be enough for a simple task whose effects we can observe and reverse. If a decision depends on conditions we may have missed, we need more knowledge of the field. A detailed answer does not reduce that need.

We can also narrow the job we give AI, rather than consult an expert for every question. We can use it to learn the terminology, identify options, or prepare for a conversation without handing over the decision. When we cannot assess the assumptions behind an important decision, getting help from someone who can may be a necessary part of the work.

## Take the same care outside your own field

When reviewing AI's work in our own field, we don't assume that a good-looking result is correct. We want to know which decisions sit behind it. Keeping that curiosity in an unfamiliar field doesn't require us to pretend we are experts. It requires us to account for the limits of what we know.

Before acting on an AI answer about an unfamiliar subject, we can ask ourselves:

**If this answer contained the equivalent of a mistake I often catch in AI's work in my own profession, would I recognise it?**

If the answer is no, “It looked right to me” remains a weak basis for a decision. We can first work out what we can check independently, what we can test on a small scale, and which part needs a review by someone who knows the field.

---

Language: English
License: CC BY 4.0
License URL: https://creativecommons.org/licenses/by/4.0/
Scope: Evren Bal-authored text, unless this article expressly states otherwise.
Excluded: Third-party material, quoted excerpts, logos, and separately marked images retain their own rights.
Attribution: Credit Evren Bal, link to the canonical source and license, and indicate changes.
Source: https://evrenbal.com/is-ai-only-bad-at-your-job
