A machine goes down at 2 a.m. The line is stopped, the shift
supervisor is standing behind you, and the fault message on the screen
is vague or absent. This is the moment the whole book is written for.
Before you touch a single wire, the most valuable tool you have is not
your multimeter — it is a disciplined way of thinking. Good
troubleshooters are not the people who memorize the most fault codes;
they are the people who follow a repeatable process even when they are
tired and under pressure.

The central idea is simple: a symptom is not a cause. A conveyor that
will not start is a symptom. The cause might be a tripped motor
overload, a failed start pushbutton, a broken wire to a PLC input, a
safety relay that has dropped out, or a rung of logic waiting on a
permissive that never arrives. Novices fixate on the symptom and start
swapping parts. Experienced technicians treat the symptom as the
beginning of a search and narrow that search methodically until only one
cause remains.

Safety comes before curiosity

Nothing in this book matters more than this section. PLC-controlled
machines move, heat, cut, and energize on their own schedule. An output
you are probing can turn on mid-scan. A drive can command a motor the
instant a permissive clears. Before you open a panel or place a meter
probe on a live terminal, you must know the energy sources present, how
to isolate them, and whether the task requires full lockout/tagout
(LOTO) or can be done safely on a live circuit with the correct rated
equipment and permits your site requires.

  • Identify every energy source: line voltage, control voltage,
    pneumatic, hydraulic, stored energy in capacitors and springs, gravity
    on raised loads.

  • Follow your site’s LOTO procedure. If you are testing live
    because the fault only appears when energized, use properly rated PPE
    and meters, and never work alone on live equipment.

  • Assume any output can turn on. Before forcing an output or
    clearing a fault, confirm nobody is in the machine and no motion will
    cause harm.

  • Prove your meter on a known live source before and after a
    zero-energy check — the ‘live–dead–live’ test.

FIELD RULE

If you would not want your reasoning read aloud at an incident
review, stop and rethink it. Slowing down for thirty seconds has saved
more fingers than any interlock.

The mindset in one sentence

Observe precisely, form one testable hypothesis at a time, and let
measurements — not assumptions — eliminate possibilities until the cause
is cornered. Everything else in this book is machinery in service of
that sentence.

Why process beats intuition

There is a seductive myth in maintenance work: that the best
troubleshooters have a kind of sixth sense, walking up to a machine and
knowing what is wrong. What actually distinguishes them is less
glamorous and far more learnable. They have internalized a process so
thoroughly that it runs in the background while they work, and that
process protects them from the two failures that sink everyone else —
jumping to conclusions and getting lost. When you watch a veteran solve
a fault quickly, you are usually watching a process executed so smoothly
it looks like instinct.

Consider what happens without a process. A pump will not start.
Someone remembers that the last time a pump would not start it was a bad
contactor, so they replace the contactor. It does not help. They replace
the start button. Still nothing. Three hours and two parts later, they
discover a blown control fuse that a two-minute check would have found
first. The parts were not wrong guesses exactly — each had failed before
— but guessing by memory of past faults is not troubleshooting. It is
gambling with the plant’s time and the parts budget.

The disciplined alternative costs a little pride and saves enormous
time. You resist the urge to fix anything until you have narrowed the
cause, and you narrow it by asking questions the machine can answer with
measurements. The pump will not start: is the controller commanding it
to start? If you can answer that in thirty seconds by going online, you
have already split the problem in half before touching a single
component.

The cost of assumptions

Every assumption you make without checking is a place the fault can
hide. ‘The sensor must be fine, it is new’ — new components fail, and a
new component installed wrong fails immediately. ‘There is power, the
light is on’ — the light might be on a different circuit than the one
that died. ‘It cannot be the program, nobody changed it’ — perhaps, but
a fault that only appears now might depend on a condition that only
arose now. None of these assumptions are stupid; they are reasonable,
and that is exactly why they are dangerous. The discipline is not to
distrust everything, but to notice when you are assuming rather than
knowing, and to convert the important assumptions into quick checks.

A USEFUL QUESTION

When you catch yourself about to replace a part, ask: ‘What have I
actually measured that proves this part is the cause?’ If the honest
answer is ‘nothing yet,’ you are guessing. Sometimes an educated guess
is worth it — but know that is what you are doing.

Symptom, cause, and the space between

It is worth being precise about three words that get used loosely.
The symptom is what you observe: the machine stopped, the reading is
wrong, the alarm is showing. The cause is the specific thing that, if
fixed, makes the symptom go away: the broken wire, the failed sensor,
the tripped overload. Between them lies a chain of intermediate
conditions, and troubleshooting is the act of walking that chain from
symptom to cause. The alarm ‘low pressure’ is a symptom; the cause might
be a failed pump, but between them lie the pressure sensor, its wiring,
its analog input, the scaling, and the logic threshold — any of which
could produce the same symptom with the pump perfectly healthy.

This is why ‘what changed?’ is one of the most powerful questions you
can ask. Machines that ran yesterday and fail today usually fail because
something changed: a part was replaced, a setting was adjusted, the
weather turned cold, a new product started running, maintenance was
performed nearby. The change is often the thread that leads straight to
the cause. When you arrive at a fault, before anything else, ask what is
different now from when it worked.

The four questions that start every job

Before you touch anything, four questions orient you and prevent most
wasted effort. What is the machine actually doing versus what it should
be doing? This forces a precise symptom instead of a vague complaint.
What changed since it last worked? This surfaces the thread that so
often leads straight to the cause. What does the machine itself report —
any alarm, any fault, any indicator? This lets the equipment tell you
what it saw. And what is the safest way to investigate this? This keeps
you and everyone around the machine intact. Ask these four every time,
and you begin every job from a position of understanding rather than
reaction. They cost under a minute and repeatedly save the hour that
panic would have burned.

Notice that none of these questions involve a tool or a component.
They are about establishing reality before acting on it. The pull toward
action under pressure is strong — a stopped line generates a powerful
urge to do something visible immediately — but the technician who pauses
to answer these four questions almost always finishes first. Speed in
troubleshooting comes from not going down wrong paths, and wrong paths
come from acting before understanding.

Confirmation bias at the panel

One trap deserves special mention because even experienced people
fall into it: confirmation bias, the tendency to notice evidence that
supports your current theory and dismiss evidence against it. You decide
early that the sensor is bad, and every ambiguous reading now looks like
a bad sensor while the perfectly good measurements that contradict you
get explained away. The antidote is to state your hypothesis explicitly
and then actively try to disprove it rather than confirm it. If you
think the sensor is bad, the strongest test is one that would prove it
good — inject a known signal, watch its own LED respond correctly. When
you hunt for evidence against your theory and fail to find any, only
then have you really confirmed it. Theories that survive an honest
attempt to break them are the ones worth acting on.

Working under pressure without losing the thread

The conditions under which you troubleshoot are rarely calm. A line
is down, every minute has a cost someone is counting, and often there is
an audience — a supervisor, an operator, sometimes a manager — watching
and occasionally offering theories. This pressure is itself a hazard to
good troubleshooting, because it pushes toward exactly the behaviors
that fail: acting before understanding, grabbing the first plausible
part, abandoning the method to look busy. Recognizing pressure as a
distinct challenge, separate from the technical fault, is the first step
in managing it. The technical fault is solved by method; the pressure is
managed by trusting the method even when it feels too slow. The paradox
worth internalizing is that under pressure the disciplined approach is
not only more reliable but usually faster, because it does not squander
time on wrong paths. The calm that looks like confidence to the watching
supervisor comes from knowing the method will get there.

There is also a communication skill that pairs with the technical
one. Keeping the people around you informed — ‘I’m checking whether the
controller is even commanding this to run, give me two minutes’ — buys
you the space to work the method rather than being pushed into hasty
action. It converts the audience from a source of pressure into people
who understand what you are doing and why. Silence invites interruption
and second-guessing; a brief narration of your reasoning earns you the
room to reason. This is not about performance; it is about protecting
the conditions under which careful troubleshooting is possible, which is
part of the job in any real plant.

The humility that makes you faster

A final element of the mindset is a specific kind of humility: the
willingness to be wrong quickly and cheaply rather than defending a
theory expensively. Good troubleshooters form hypotheses eagerly but
hold them loosely, ready to discard one the moment a measurement
contradicts it. The failure mode is emotional attachment to a theory —
having decided it is the sensor, you keep finding reasons it must still
be the sensor as contrary evidence piles up. The discipline is to treat
each hypothesis as disposable, valuable only until a measurement kills
it, at which point you drop it without regret and form the next.
Paradoxically, the technician most willing to be wrong is the one who
reaches the right answer fastest, because they waste no time defending
dead theories. Being attached to finding the truth, rather than to being
right, is what keeps the search moving.

Leave a Reply

Your email address will not be published. Required fields are marked *