Most detection programs are reactive by design. An alert fires, an analyst investigates, a ticket closes. That loop catches what your tooling already knows how to catch — and quietly misses everything it does not. Threat hunting exists to close that gap. It is the deliberate, human-led search for adversary activity that has evaded automated controls, conducted before an incident is declared rather than after. Done well, it surfaces the dwell-time attacker who has been living off the land for weeks. Done poorly, it becomes a directionless scroll through logs that produces a warm feeling and no findings.
This guide lays out a practical way to run a hunt: what hunting actually is, how to build a hypothesis, where MITRE ATT&CK fits, which data you need, how to gauge your team’s maturity, and the errors that sink otherwise capable programs. None of it requires a dedicated threat-hunting platform. It requires telemetry, a question, and discipline.
What is threat hunting, really?
Threat hunting is the proactive search for signs of compromise that existing detections did not flag. The operative word is proactive. If a SIEM rule generated the lead, that is alert triage, not hunting. If an EDR quarantined the file, that is prevention. Hunting begins where automation ends — with the assumption that a capable adversary is already inside, and that their activity is hidden inside normal-looking telemetry.
That assumption matters. A hunt is not a vulnerability scan and it is not a compliance audit. You are not looking for missing patches or misconfigured buckets; you are looking for behavior. Did a service account authenticate from a workstation it has never touched? Did rundll32 spawn a network connection to a rarely seen destination? Did an encoded PowerShell command run outside a maintenance window? These are questions about what actors do, not what your environment is.
A useful distinction: hunting produces two kinds of value. The obvious one is finding the intruder. The less obvious — and more common — one is producing new detections. Every hunt that turns up a behavior worth watching should end with a durable detection rule so the same behavior never requires a manual hunt again. Over time, that feedback loop is what separates a program that scales from one that depends on a single gifted analyst.
How do you build a hunting hypothesis?
Hypothesis-driven hunting is the discipline that keeps a hunt from wandering. Instead of “let’s look at the logs,” you start with a falsifiable statement: An adversary is using scheduled tasks for persistence on our domain controllers. That sentence tells you exactly what data to pull, what normal looks like, and what would confirm or refute it.
Good hypotheses come from three sources. First, threat intelligence: a report on APT groups or other actors targeting your sector gives you specific, patient techniques to look for — precisely the kind of quiet, persistence-oriented tradecraft that automated detections are least likely to catch on their own. Second, your own environment: knowledge of what is unusual for your network — which accounts should never log in interactively, which hosts should never talk to the internet. Third, the ATT&CK matrix itself, worked technique by technique, asking “could this happen here, and would we see it?”
A workable hypothesis has three parts. It names a behavior (credential dumping, DLL side-loading, data staged in an archive). It names the data that would reveal it (process creation logs, LSASS access events, file-creation telemetry). And it defines the outcome — what you expect normal to look like, so that deviation stands out. Write it down before you query anything. If you cannot state in advance what a positive result looks like, you are not hunting; you are browsing.
Where does MITRE ATT&CK fit?
ATT&CK is the common language of modern hunting. It catalogs adversary tactics (the why — persistence, lateral movement, exfiltration) and techniques (the how — the specific methods used to achieve each tactic), all grounded in observed real-world behavior. For a hunter it does three things.
It structures coverage. Mapping your existing detections to the matrix reveals blind spots — tactics where you have deep visibility and tactics where you have none. Those gaps become your hunting backlog. It seeds hypotheses. Each technique page describes how the technique is carried out and, critically, which data sources reveal it, so you can move from “I should check for lateral movement” to a concrete, testable query. And it standardizes communication. When a hunter writes up a finding as “T1053.005, scheduled task persistence,” every downstream reader — detection engineer, IR lead, CISO — understands it immediately.
A caution: coverage of the matrix is not the same as security. It is tempting to chase a green cell for every technique, but some techniques are rare, environment-specific, or better handled by prevention. Use ATT&CK to prioritize by what adversaries relevant to you actually do, not to fill in a scorecard. The matrix is a map, not the territory.
What telemetry does hunting require?
You cannot hunt for what you cannot see. The single largest determinant of a hunt’s success is the quality and retention of your telemetry. At minimum, a credible program needs visibility into a few core domains.
Endpoint is the richest. Process creation with full command lines, parent-child process relationships, module loads, and network connections originating from processes are the backbone of most hunts. Command-line arguments in particular are where obfuscation, living-off-the-land binaries, and encoded payloads reveal themselves. Identity and authentication telemetry — logon events, Kerberos activity, privilege changes, and directory modifications — is where lateral movement and credential abuse surface. Network data, even summarized flow records or DNS query logs, catches beaconing and exfiltration patterns that never touch an instrumented endpoint. And cloud and SaaS audit logs are increasingly non-negotiable as workloads move off traditional infrastructure.
Two practical constraints shape everything. The first is retention. If your logs roll off in seven days but adversary dwell time runs into weeks or months, you cannot hunt the past — you can only hunt the present. Aim for retention that outlasts realistic dwell time. The second is normalization. Telemetry scattered across incompatible formats forces the hunter to spend their time as a data engineer. NIST’s guidance on log management, NIST SP 800-92, remains a sound reference for building the collection and retention foundation that hunting depends on.
How mature is your hunting program?
Maturity models help teams see where they are and what “better” looks like. A serviceable four-stage model runs from ad hoc to leading.
At the initial stage, hunting happens occasionally, driven by one curious analyst, with no repeatable process and little data. Findings, if any, do not turn into detections. At the minimal stage, the team consumes threat intelligence and runs searches based on indicators of compromise, but the work is still largely reactive to external reports. At the procedural stage — where most competent teams should aim to live — hunts follow documented, repeatable procedures, draw on rich telemetry, and are mapped to ATT&CK. Every hunt is logged and every useful finding becomes a detection. At the leading stage, much of the procedural work is automated, freeing hunters to develop novel hypotheses, and the program continuously feeds detection engineering at scale.
Two things move a team up this ladder: better data and better process. Tooling helps, but a team with strong telemetry and disciplined hypotheses on modest tools will out-hunt a team with an expensive platform and no method. Assess honestly. Most organizations overestimate their stage because they conflate having an EDR with having a hunting practice.
Hunting also does not live in isolation. It feeds, and is fed by, the broader security operations function and the controls protecting your data security posture. A finding from a hunt should tighten a detection, inform an access-control decision, or reshape what you collect.
What mistakes derail hunts?
The most common failure is hunting without a hypothesis. Open-ended log-diving feels productive and almost never is; it exhausts the analyst and produces nothing durable. Start with a question you can answer.
The second is neglecting the write-up. A hunt that finds nothing is still valuable — it establishes what normal looks like and rules out a threat — but only if it is documented, so the next hunter does not repeat it and so a null result today can be compared against a suspicious result tomorrow. A hunt with no artifact is a hunt that never happened.
The third is failing to operationalize findings. If a hunt reveals a detectable behavior and no detection rule follows, you have committed to hunting that same thing manually forever. The fourth is chasing tools over telemetry: buying a platform before you have the data to feed it. The fifth is measuring the wrong things — counting hunts run or hours spent instead of detections created and blind spots closed. The point of hunting is not activity; it is coverage that did not exist before.
Frequently Asked Questions
How is threat hunting different from incident response?
Threat hunting is proactive and happens before an incident is declared — you are searching for compromise that automated controls missed, without a confirmed alert to react to. Incident response is triggered after a confirmed or suspected incident and focuses on containment, eradication, and recovery. A hunt that finds something real hands off to incident response.
Do we need a dedicated tool to start threat hunting?
No. The prerequisites are good telemetry — especially endpoint process and command-line data — sufficient retention, and a query interface you can search it with. A capable analyst with a hypothesis and access to well-normalized logs will out-hunt a team that bought a platform but lacks the underlying data. Invest in data and process before tooling.
How does MITRE ATT&CK help a hunting program?
ATT&CK provides a shared catalog of adversary tactics and techniques grounded in real-world observation. Hunters use it to map existing detection coverage and expose blind spots, to generate concrete testable hypotheses from technique descriptions, and to communicate findings in a standard vocabulary that detection engineers and leadership immediately understand.
How often should we hunt?
There is no universal cadence, but hunting should be scheduled and recurring rather than opportunistic. Many procedural-stage teams run structured hunts on a regular rhythm, prioritized by threat intelligence relevant to their sector and by ATT&CK coverage gaps. Consistency matters more than frequency: a documented monthly hunt beats sporadic marathon sessions that leave no repeatable artifacts.
What should a hunt produce if it finds nothing?
A null result is a legitimate, useful outcome. A “negative” hunt should still produce a documented record of the hypothesis, the data queried, and what normal looked like. That baseline lets future hunters compare against it, prevents duplicated effort, and sometimes reveals a visibility gap — you may find you could not have detected the behavior even if it were present.
