guide
33 min read
FreeHow to Read and Validate ClusterHawk Reports
This guide teaches analysts how to read platform-generated reports, what to focus on, and how to validate claims against the underlying artifacts.
By Chawkr Reports
15/08/2025
How to Read and Validate ClusterHawk Reports
Purpose
This guide teaches you how to read a single platform-generated report, what to focus on, and how to validate claims against the underlying artifacts. It reflects the current report format. Section numbers below match the headings the report emits.
Who the report is written for. Threat hunters. The report assumes you already have the analysis job open, with every raw score, asset listing and cluster in front of you. It does not restate that data, but it tells you what the data means, what is worth your day, and what is not. If you need something for a non-hunting audience, take the Executive Summary and leave the rest.
Two report forms. Every analysis job produces the standard analysis report described here. A pilot investigation deliverable is the same report plus three manually-appended annexes (per-IP disposition, hunting pack, method glossary).
Quick Start: What to Extract First
- Section 2.B (Cluster Portraits): read each portrait's closing tracking verdict first, then extract the anchors it marks worth following. The fingerprints and queries are still here; the verdict tells you which of them deserve your time. Three clusters are profiled per job for you.
- Section 10.2 (Intelligence Collection Value): the ranked list of what to chase, with access cost and whether each anchor is trivially discoverable anyway. If you have one hour, read this and Section 1.
- Section 7.1 (Tracked Infrastructure Families and the Anchor Register): the patterns that recur across clusters and providers, and the one place every anchor is fully defined.
- Section 6 (Anomaly & Outlier Intelligence): identify the CRITICAL and HIGH anomalies and, crucially, the Noise/Outlier Re-Analysis (6.B); the noise re-analysis surfaces structurally unique assets the main clustering misses.
- Section 9 (Detection Engineering): each rule now opens with its detection objective, including who else legitimately matches and what would defeat the rule. Read the objective before the syntax.
- Section 4 (Hypothesis Engine): read the competing hypotheses and whether either cap was applied before trusting any single attribution.
- Section 5 (Priority Actions and Hunt Procedures): what to do, ordered by expected yield.
The Cluster Portraits (Section 2.B) are still the richest section, but the tracking verdicts and the Anchor Register are what make it usable. The verdict tells you which patterns identify an operator rather than a product, and the Register gives every anchor a single definition the rest of the report refers back to. Most other sections reference these profiles by cluster number, and every anchor by register name.
Core Principles
- Cluster Portraits (Section 2.B) are the crown jewels: these carry the argument about what each cluster is, what it is worth, and whether it can be followed.
- The profile is the signature, not the query. This one might confuse people, so it has its own section below.
- Read the noise, not just the clusters: the Noise/Outlier Re-Analysis (6.B) routinely isolates the single most interesting asset in a dataset. A pivot never reaches it; the re-analysis does.
- "So What?" mindset: focus on operational impact and intelligence value, not raw lists.
- Data completeness: exact values belong in Section 11 (Technical Appendix); narrative sections summarize patterns.
- Anomalies are dataset-relative, not a verdict of "bad": an anomaly is a host that does not lean the way the rest of this dataset leans, a structural outlier against the dataset's own center, not against an absolute model of malice. Severity scales with how far the host sits from that center: scores at or above the 95th percentile are emitted as CRITICAL. Read "flagged" as "doesn't fit this population," then decide whether not fitting is interesting, which depends entirely on what you seeded (see Section 6).
- A quiet anomaly result is a real result: a job with zero or only low-severity anomalies is not a failure or an empty report; it means the dataset is structurally uniform (one template, one operator profile, tight cohort), so nothing meaningfully deviates from the bulk. The detector reports what is there; when the data is homogeneous, "nothing stands out" is the honest and correct finding.
- A short report is a real result too. The report is sized by how much there is to argue, not by how much effort went in. A small job produces a short report, and that is the correct outcome rather than a sign the analysis gave up.
- The compound is the finding: a single fingerprint match is noise; the multi-parameter profile is the signature. Count evidence classes, not values, when you judge how strong a compound is.
- Absence is evidence, and sometimes it is the best evidence you have. What a cohort refuses to expose is a distinction you can hunt on, provided the rest of the population does expose it.
- Ask which layer an anchor sits in: a composite that fingerprints a vendor's firmware will match beautifully on thousands of clean devices of the same model. Durability alone does not make something trackable. Durable, selective, and operator-side does.
- The verdicts are the product: every profiled cluster gets a track / do not track call with a falsification condition attached, and every anchor is defined once in the Register. If you take three things out of a report, take those and the ranked collection list in 10.2.
The Profile Is the Signature, the Query Is a Projection
This is the most consequential thing to understand about the report, and the easiest to get backwards.
A cluster's profile is everything the analysis established about it: the features that define it, the features it lacks, products, certificate conventions, transport fingerprints, naming.
The generated query is that profile projected onto the fields scan data provider happens to expose as a supported filter. Scan data providers support a fixed vocabulary of filters. A defining feature with no filter mapping gets dropped from the query string. It stays in the profile and is still real, and it is frequently still searchable as free text even though no official filter exists for it.
Three consequences worth holding on to:
- A sort query does not mean a small cluster. A cohort can be highly distinctive and still emit a weak query, purely because its strongest features are unsupported fields. The report says which of the two it is assessing, and you should check that it did.
- A match count measures the query, never the cluster. It tells you how that conjunction behaves against one provider's index. It does not tell you how recognisable the infrastructure is.
- Where no query was generated at all, that is a statement about filter coverage. It is not a statement that the cluster cannot be re-identified.
Terms beginning with - are exclusion clauses. They come from what the cohort does not carry. Because providers can express "this field does not exist," the query negates the value that hosts outside the cluster
most commonly carry instead. An exclusion only survives into the query when the surrounding population carries that
field substantially, so every negation you see is a real distinction rather than padding. On cohorts defined by what
they lack, the exclusions are usually the terms doing the selective work, and they are also how you trim a broad
pattern down to something you can actually work with.
Reading Order and Focus Areas
1. Executive Summary & Recommendations
- 1.1 Executive Summary: BLUF, Infrastructure Overview Table, Key Findings, Risk Assessment.
- The standing signature qualifier lives here, stated once for the whole report: every query, pivot and detection signature in the document is a candidate pattern. Matches identify hosts sharing that pattern, they require per-host verification, and they are not a confirmed-indicator list. That qualifier holds for everything downstream and is deliberately not repeated under each rule.
- 1.2 Priority Actions Reference: the top actions, drawn from Section 5.
- 1.3 Attribution Summary: the primary hypothesis with its rubric percentage, and whether a non-discrimination cap was applied (references Section 4). If the verdict is capped, treat attribution as undetermined.
2.A Infrastructure-Cluster Matrix (Section 2.A)
- Cluster groupings by core technology and possible operational role.
- Hunt Priority and the maturity assessment per group.
- MITRE ATT&CK tactics per group, each carrying a side: whether the tactic describes something the hosts are equipped to do, or something an attacker would do against them.
- The Infrastructure Inventory Summary (asset counts and primary providers per cluster).
2.B Comprehensive Cluster Portraits (Section 2.B)
The three most significant clusters are profiled in full. Selection is deterministic rather than a matter of taste: detection signature first, then flagged anomalies, then whether the cluster yields a candidate anchor, then confirmed known-exploited vulnerabilities, then size.
Each portrait runs as four numbered subsections, then a hosting and trackability block, finally the verdict.
1. Threat Intelligence & Technology Stack
- Detection Signatures (labels and their confidence splits, for example
CobaltStrike:50%/GoPhish:50%). Read what the signature names. The rule set covers malicious tooling and also device classes, service roles, security tools and legitimate applications, so a match may be saying "this is a Cobalt Strike server" or "this is a mail server" or "this is a Plex box." Those are different findings and the report treats them differently. A match is also a lead to confirm rather than a fact to report: the report checks it against the host's own banner, certificate and service surface, and says whether that evidence backs it up. - Product and technology families, and whether the version spread indicates maintenance or its absence.
- Vulnerability Intelligence: whether a CVE fingerprints the product (incidental, every host on that version carries it) or the operator's configuration posture (a signature). The identifiers, counts and KEV status live in 11.3.
2. Infrastructure & Network Profile
- Network services: the exposures that carry meaning, including a service the host's role does not need, a critical exposure, or a surface whose absence is itself the finding.
- Reverse-proxy and fronting posture: whether anything refuses direct requests without an onward redirect.
- Certificate and cryptography: issuer conventions read through durability and residue, not restated as statistics.
- Industrial protocols (ICS): BACnet, Modbus, Siemens S7 and similar, when the assets expose OT services. A high-signal finding that usually sits outside standard IT detection tooling.
- Exposed databases: MongoDB, Redis, Elasticsearch, MySQL, and whether authentication is enabled.
3. Cluster Validation & Uniqueness
- The compound query that projects the cluster's profile, with its real-world match count.
- Alongside it, the reduced pivot: a three to five term subset, durable and selective enough to actually return a population. The full conjunction is precise and usually unrunnable. The pivot is what you run first; you narrow to the full conjunction once you have hosts to narrow. Pivot terms may be positive or negated, and a negation counts as a full term.
- A commodity match count (millions) flags a generic cluster. Baseline it, do not track it as an actor.
- Where the emitted query understates the profile because the provider cannot express the cluster's strongest features, the report says so in a clause. Read that clause. It is the difference between "we cannot find these hosts" and "Provider has no filter for the thing that identifies them."
Read a zero-match result against the query's specificity, not as a uniqueness score. This is the single most common misreading of the report. A conjunction stacking fifteen or more exact terms, including content hashes and certificate fingerprints, returns zero for almost any host group, benign ones included. Zero there is the expected outcome and tells you nothing about whether the infrastructure is unusual, purpose-built or targeted.
Zero on a low-specificity query is different: port plus product plus organization, no stacked hashes, no matches means something. That is a real uniqueness signal.
What a zero-match compound query does give you either way is a precise re-identification signature: a future host matching the whole conjunction is a strong candidate for the same cohort. Candidate, not confirmed. Verify per host before you act, and never publish these as an IOC feed or wire them to automatic blocking.
- If a cluster's profile contains no term the analysis can argue is non-base-rate, the report says so and emits no reduced pivot for it. That absence is deliberate and it is information: a pivot built from a default SSH algorithm list or a common-stack TLS fingerprint returns the base rate, which is worse than nothing because it looks like a result.
4. Statistical Analysis & Overlap
- Shared components and fingerprints with other clusters, and whether the overlap is a genuine relationship or a base-rate coincidence. The report lists some overlaps specifically in order to dismiss them.
- The adversary-versus-victim verdict, which is one of the two load-bearing judgements in the portrait. Everything downstream inherits it: the tactics cell in 2.A, the technique readings in Section 3, the role classification in 7.2, and how any Section 9 rule scoped to that cluster is framed.
- The medoid (most representative member) and borderline IP: what defines the cluster's identity and where it blurs.
Hosting, Identity and Trackability (closes each portrait)
- Hosting strategy: provider spread, geography, and what the procurement shape implies. The percentages themselves are in 2.A and 11.3.
- Cluster identity: the highest-importance differentiating features, including defining absences, which are part of the identity rather than a separate list of things not observed.
- Instance validation: the medoid's defining traits and how the borderline asset diverges.
- Quality, but only where it changes the verdict. A cluster whose members the methods grouped confidently while the hosts themselves disagree about what they are cannot be tracked as one thing, and the portrait says so. Where quality has no bearing on trackability, the portrait skips it and the numbers stay in Section 12 and 11.3.
Tracking verdict (the last line of every portrait)
One line, and it is the line to read if you read nothing else in the portrait:
Tracking verdict: Track / Track with verification / Do not track, with the anchor's name, its layer, and what would falsify it.
- Layer answers whether the anchor identifies the operator's own configuration, the vendor's product, or neither. A composite that fingerprints a camera model will match cleanly on thousands of unrelated units of the same model. It is an inventory finding, not a campaign anchor, and the report will say so rather than letting a clean match flatter you.
- Falsified by is the result that should make you drop the anchor. Read it before you commit a day to a lead. If it says "none stated," the anchor is not ready to track.
- Do not track is a real verdict, not a failure. Most clusters in most jobs earn it. When a cluster earns it, the portrait still carries the full evidence, because the profile is what justifies the negative call.
Where the numbers live. The portraits carry the argument. The reference tables (ports, products, CVEs, certificate statistics, provider percentages) live once, in Section 11.3. If you want the figures, go to the appendix; if you want to know what they mean, stay in the portrait.
Why this section enables actor tracking:
- Fingerprints and deployment templates persist across IP rotation, so you can follow infrastructure whose addresses churn.
- Compound profiles become predictive alerts. The compound, not any single parameter, is what makes the alert high-precision.
- High Metric Amber differentiators identify signature behaviors: the more a feature is over-represented against the dataset, the more distinctive it is.
Pro tip: extract the fingerprints and the reduced pivot from each portrait, put the pivot into continuous monitoring, and pull exact values from Section 11 for platform integration. Never deploy a single-parameter match as a detection. It will be noise.
3. MITRE Mapping (Section 3)
The section opens with the modelling basis and its limits, and that paragraph is worth more than the table under it. The analysis models observable configuration: what a host is equipped to do, and what its setup implies about who assembled it. It cannot model traffic, payload, victim, authentication or sequence. That rules out kill-chain reconstruction, dwell time, intent, and any claim that an observed capability was ever exercised. A technique mapped here says the capability is present. It never says it was used.
Every row declares which side it is read from. A technique is either an adversary capability the host is equipped to exercise, or a victim exposure that makes the host a target. The same open port supports both readings, and a table that does not say which is quietly asserting the more alarming one. The reading also has to match that cluster's adversary-versus-victim verdict in its portrait, so a cluster called victim-side exposure cannot carry an adversary-capability row.
That constraint exists because the failure it prevents is a real one: mapping Command and Control onto clusters the analysis has just concluded are ordinary exposed web servers tells anyone skimming the table that the job found C2 infrastructure, when it found the opposite.
Progression narratives across clusters (staging to relay to C2) appear only when a coordination signal supports them. Absent that, the report says so rather than assembling a story out of host typology.
4. Hypothesis Engine (Section 4)
- The competing hypotheses, typically a null (H0: unrelated or typological), and two to three alternatives (multi-operator collection, coordinated service, red team versus malicious, IAB or APT staging).
- For each: supporting evidence with cluster references, disconfirming indicators, and a scoring table across the five summed dimensions (Feature Distinctiveness, LIME Coherence, Detection-Signature Consistency, Query Validation, Cross-Cluster & Anomaly Consistency). The scoring is a table because it is scoring; the prose underneath it carries the disconfirming indicators, which is where the analytical value sits.
- The score is a rubric percentage, not a calibrated real-world probability, and the report says so explicitly. Read it as relative ranking, not odds.
Agreement between the report's own methods is not corroboration. Feature importance, instance-level explanation and cross-cluster consistency are three ways of looking at one input. When they line up, that means the finding is not an artefact of a single method. It does not mean three sources confirmed it. The report states this once and caps the confidence accordingly. Independent corroboration means something from outside the dataset: registration or passive-resolution history, a different scan provider, published reporting.
Two caps, applied in a fixed order. The cluster-quality cap comes first: it is a property of the evidence, and where every cluster central to a hypothesis sits in the weak stability band, the hypothesis cannot be reported above Medium-High no matter what its tiers total. The non-discrimination cap comes second, on the already-capped figures, because it is a property of the comparison between two hypotheses. If both bind, the report says so and gives the order.
The Non-Discrimination Cap is the most important thing not to misread. When two hypotheses score within about 5% of each other, the engine caps them and states that the data cannot discriminate between them. A 75% / 73% split does not mean "two close but distinct possibilities": it means do not pick a winner; the evidence to separate them (victimology, payload, temporal batching) is not in passive metadata. Honor the cap, keep both hypotheses active and drive targeted collection.
Confidence interpretation for the uncapped spread: 85-95% High, 70-84% Medium-High, 50-69% Medium, 30-49% Low-Medium, 10-29% Low. Note the ceiling: the rubric does not emit 100%, and on a job where every generated query is high-specificity it will rarely reach the top band at all, because Query Validation contributes nothing in that case.
5. Priority Actions and Hunt Procedures (Section 5)
All actionable recommendations in the report live here, and they are hunting actions rather than remediation advice. You do not own these assets and the report does not pretend otherwise.
- Priority Actions, ordered CRITICAL / HIGH / MEDIUM, each naming the cluster it applies to.
- Hunt Procedures, ordered by expected yield: what to enumerate, on which assets, and what result would change the reading. Anomaly investigation targets and pivot queries are here, referencing Section 9 by name.
- Pivot Indicator List: the addresses whose signature names malicious tooling or a malware, C2, botnet, stealer or RAT family. Device classes, service roles, security tools and legitimate applications are excluded by design, and anything the report cannot classify from the signature name is excluded and flagged for analyst review. On a job with no tooling detections there is no list at all, and the report says so rather than handing you a list of your own mail servers.
- Collection Priorities: ordered actions matching the Section 10.2 ranking.
The pivot list is not a blocklist. These are starting points for investigation, not confirmed indicators, and the report will not issue a blocklist even when it could. Submitted datasets routinely contain live business web properties, payment service domains and commercial marketplaces. Blocking on cluster membership blocks those too.
6. Anomaly & Outlier Intelligence (Section 6)
Read all of it: the most interesting asset in a dataset usually lives in 6.B, not 6.A.
Read every anomaly against the dataset you submitted. The detector finds hosts that deviate from this dataset's structural center; it does not score against a fixed model of "malicious." That makes interpretation entirely dependent on what you seeded:
- On a confirmed-bad or C2 seed, the anomalies are the hosts that break the malicious bulk, which can mean benign contaminants, a misrouted host, or an operator who is structurally different from the rest. "Flagged" here does not mean "the worst one"; it means "the one least like the others."
- On a mixed or broad seed, the anomalies are the genuine structural oddities: the bespoke host in a sea of commodity tooling, the standalone node deliberately separated from the fleet.
Same host, different dataset, different verdict. That is correct behavior, not instability. The detector tells you what is structurally different; you decide whether different is interesting. This is also why the verify-ground-truth step matters: confirm a flagged host against its own banner and certificate before deciding what its difference means.
A low-anomaly or zero-anomaly result is a valid finding. If a job returns few or no anomalies, the dataset is structurally uniform: a tight, templated population where nothing meaningfully stands out. That is a real, honest answer (often a strong signal of single-template infrastructure), not an empty report.
6.A Main Clustering Anomaly Analysis
- IPs with unstable cluster membership across the clustering runs (assigned to many distinct clusters, low neighbor consistency, high neighbor turnover), each with a severity from CRITICAL down to LOW.
- The Top-N most anomalous IPs with their distinct-cluster counts and key instability pattern.
- Common anomaly patterns and which clusters the instability concentrates in. Where the anomaly load clusters into one or two groups rather than spreading, that concentration is itself a finding about where the cluster boundaries are weakest.
6.B Noise/Outlier Re-Analysis: the high-value move
- The noise partition (Cluster -1) is re-analyzed on its own, producing noise sub-clusters plus residual outliers.
- Method Vela flags the structurally anomalous noise IPs, with a percentile and a nearest-main-cluster pointer.
- This is where deliberately-separated infrastructure surfaces, for example a pure single-tool detection sitting outside every main cluster. A same-indicator pivot never reaches these; the re-analysis does.
- Cross-system signal: an IP flagged in both 6.A and 6.B is the highest-priority target in the dataset.
6.C Anomaly Intelligence Categorization Matrix: anomaly type, detection source, indicators, threat potential and recommended action.
6.D Anomaly Cluster Profiles: per-IP narrative, what defines its assigned cluster and exactly how it diverges.
6.E Tool-Detected Outlier Profiles: assets carrying offensive-tooling labels that landed in noise, with their similarity-group and reverse-similarity-group memberships (low-importance features shared with other noise IPs, an infrastructure-reuse signal).
7. Attribution, Operational Intelligence & Victim Analysis (Section 7)
7.1 Attribution Indicators Analysis
- Infrastructure behavioral patterns and the attribution challenges they raise.
- Technical Tradecraft Assessment: default versus custom configurations, OPSEC sophistication, fronting, naming discipline.
- Procurement and selection patterns: provider strategy, regional alignment, certificate procurement.
Tracked Infrastructure Families. A family is a pattern that recurs across two or more independent contexts, where a context can be a cluster, a provider, a platform or a device family, on an indicator that survives base-rate discipline and sits outside the device-model layer. That last condition does the real work. Without it a vendor certificate convention spanning two device clusters and two carriers would qualify while an operator's build standard recurring across bastions, print controllers and database hosts inside one cluster would not, which is exactly backwards. Commodity provider overlap, common-stack TLS fingerprints and default algorithm lists never qualify.
An empty families table is a finding, not a gap. When nothing clears the gate the report says so and names what came closest and what it lacked. That tells you where to collect next.
The Anchor Register. Every anchor the report names anywhere, whether in a portrait verdict, a family, a detection signature or a collection target, is defined once here: name, parameters, the evidence classes those parameters span, layer, verdict, reduced pivot, and what would falsify it. Everything else in the report refers to it by name and carries none of the detail.
Names are derived from what the anchor is, not from the cluster number it landed in, because cluster ids are assigned per job and mean nothing outside it. Use the register name when you record an anchor in your own tooling.
7.2 Infrastructure Role and Vulnerability Analysis
Roles are assigned by model, not by pattern match. For each role the report states what the role predicts a host would present, what was actually observed, and what is missing or what competing role explains the same observables equally well. That third part is what makes the classification falsifiable, and where the prediction fails the honest output is Undetermined with the failure named. "Port 443 is open, therefore C2" fits most of the internet, and the report will not write it.
Operator roles (Command and Control, Data Staging, Proxy/Relay) cannot be assigned to a cluster the portrait verdicted victim-side. For those the report describes what the host is in its own right, and where relevant what it would be useful for if someone recruited it.
Critical exposure identification names which CVEs, services and configurations constitute the exposure per cluster. It does not rank them, because Section 8 ranks visibility gaps and Section 10.2 ranks collection value, and a third ranking of the same clusters under a different heading would be duplication.
8. Defensive Intelligence & Gap Analysis (Section 8)
Gaps here are hunting gaps, not remediation gaps.
A gap is a detection objective this dataset cannot support a rule for. The report names the objective, why the artefacts do not reach it, and the specific observable that would close it. "Cluster 2's content hashes are fully fragmented, so the objective of re-identifying the cohort by content cannot be met; a stable certificate or service-stack composite would be needed" is a gap. "Cluster 2 has poor coverage" is not, and you should push back if you see it.
The section carries the job's single ranked list of visibility gaps, covering detection blind spots, fingerprint fragmentation that defeats transport-layer detection, certificate conventions outside any CA trust chain, and asset classes that produce no usable telemetry at all.
9. Operational Threat Hunting & Detection Engineering (Section 9)
Every rule opens with a detection objective. A rule without one is a query, not a detection, and the objective is where most of the usable information sits:
- Objective: what the rule is meant to surface, in behavioural or structural terms rather than a restatement of its own syntax.
- Layer: whether a match identifies operator configuration, a device model or vendor firmware, or is undetermined. This matches the layer recorded in the Anchor Register.
- Expected match population: who else legitimately matches. Every rule here has a false-positive population, and naming it is what lets you triage. If the report cannot describe who else matches, it says so rather than pretending the rule is clean.
- Defeated by: the cheapest change that makes the rule stop working. A content hash dies when the page changes, a certificate convention dies when the operator regenerates, an absence-based rule dies the moment the cohort exposes the service. This is what tells you how long the detection is good for.
Where a cluster warrants no rule, the report states the objective anyway and explains why the data cannot serve it. That is the honest form of a coverage gap and it is more useful than silence.
Each rule also ships with its reduced pivot and its falsification condition, the latter by reference to the Anchor Register rather than repeated in full.
How rule confidence is counted. Confidence tracks the number of independent evidence classes a rule conjoins, not the number of values it stacks. The classes are service surface, product identity, certificate convention, transport fingerprint, naming, content hashing, and absence. Six hashes are one class. Three parameters drawn from three classes beat six drawn from one. Three or more classes is High, two is Medium, one is an investigation starting point however many values it lists.
This is why a short rule can outrank a long one. The strongest anchor in a report is often a three-term composite spanning an SSH configuration hash, a port and a private naming convention, while a nine-term conjunction of content hashes is one class wearing a lot of numbers. An exclusion clause is a full term and its class counts, so a rule combining a service surface, a product identity and a well-founded absence spans three classes and earns High Confidence on that basis.
Rules may draw on the whole profile, including features no scan provider exposes as a filter, because rule syntax is not bound to one provider's vocabulary. Pivots may not, because a pivot has to run.
10. Operational Security & Risk Assessment (Section 10)
10.1 Operational Security Assessment: sophistication reads per cluster, deployment template consistency, infrastructure persistence, and whether the population looks provisioned or accumulated. Watch for the firmware-versus-operator check: a rigid template riding on a single device family is the vendor's image, not a provisioning toolkit.
10.2 Risk Assessment Framework: operational risk per cluster (severity against exposure), then the report's single ranked list of tracking targets.
Intelligence Collection Value is where you decide what to do first. Every anchor in the Register appears here, ranked best first, with two columns worth reading carefully:
- Confirming source and access cost. What you need in order to confirm the anchor, and roughly what it costs you. "Obtain registration history" is a five-minute job if you already have the access and a procurement conversation if you don't. The report says which.
- Trivially discoverable. Whether the anchor would fall out of one obvious public query anyway. A wildcard certificate convention usually would. A configuration uniformity spanning heterogeneous platforms usually would not. An anchor that is both durable and non-obvious is worth more of your day than one that is merely durable, and this column is the only place the report tells you so.
The subsection closes with the negative verdict: every "do not track" collected into one statement of what is not worth pursuing in this submission, and why. Read it. A negative result you can point at is worth more than a hedge, and it is what stops the same submission being re-analysed next quarter on the assumption there was something in it.
11. Technical Appendix (Section 11)
- 11.1 Infrastructure Asset Lists: a census. Every cluster in the job gets one row (id, label, asset count, classification), ordered by size, with the reserved partitions last. Clusters that were profiled, that carry a flagged anomaly, or that are reserved partitions also get a short note underneath. Everything else is the row alone, which is what keeps this readable on jobs with a hundred clusters.
- 11.2 Cross-Cluster Correlation Matrix: the top shared-indicator overlaps with confidence.
- 11.3 Infrastructure Profiles: detailed profiles for the three profiled clusters plus any flagged anomalous cluster not already among them. Exact fingerprint values live here, along with the reference tables the portraits do not repeat: ports, products, CVEs, certificate statistics, provider percentages. This is the source for platform integration.
- The Detection Query row carries two queries: the full profile conjunction with its match-count band, and the reduced pivot beneath it. Copy the pivot when you want hosts back and the full conjunction when you want to narrow a population you already have. Exclusion terms appear exactly as generated.
- The Tracking verdict row carries the verdict word and the anchor's register name, and nothing else. The layer, the falsification clause and the reasoning are in the portrait and the Register. If you want the full verdict, that is where it lives.
12. Clustering Quality and Methodology Assessment (Section 12)
- 12.1 Dataset Composition: total IPs, clustered assets, noise and outliers (Cluster -1), honeypots (Cluster -2), NIL (Cluster -3). These counts are derived once, here, and every other section cites them unchanged. If two sections disagree on a count, that is a defect worth reporting.
- 12.2 Quality Metrics: per-cluster ratings for Metric Quartz (stability), Metric Obsidian (distinctness) and Metric Topaz (distribution), rolled into an Overall Evaluation (Very Good / Good / OK / Bad). Treat "OK" clusters as hypotheses.
- Content agreement sits beside those as a second and independent axis, measured on the hosts' own observed data rather than on the numeric matrix. The two routinely disagree, and the disagreement is the interesting part: a cluster the methods held firmly while its members agree least about what they are cannot be followed as one thing.
- Detection signature status for the job: whether anything matched, stated once. If nothing did, every attribution downstream rests on infrastructure pattern analysis alone.
- Query-specificity reading: each emitted profile classified as high or low specificity, with what its match count therefore does and does not license. This is the only place that calibration argument is made in full.
- Anomaly rate: the headline count and percentage flagged, with the severity breakdown.
Read the quality rating honestly. A high "Very Good" share means clean typological separation, not proof of a single operator. Separation is about how distinct the groups are, not who runs them.
Investigative (Pilot) Report
The pilot investigation deliverable is the standard analysis report above, delivered findings-first, with three extra annexes appended by the analyst.
- Annex A: Per-IP Disposition. The dataset composition (total, clustered, outliers -1, honeypots -2, NIL -3, usable) and a table of the notable IPs from the submitted list (any IP with a tooling detection, a CRITICAL or HIGH anomaly, or a honeypot or NIL disposition), each with its cluster, signature, anomaly severity, and disposition. A companion CSV carries every submitted IP so you can audit any single address.
- Annex B: Hunting Pack. The per-cluster compound queries with their real-world match-count band, plus the pivot indicator list. The curated detection rules stay in Section 9. Annex B is the raw, copy-paste hunting material and the match-count context.
- Annex C: Method Glossary. Plain-language meanings of the metric and method codenames (Metric Quartz and the other quality metrics, Method Vela, Metric Opal and Ruby, and the clustering-method names) plus the confidence-band caveat: qualitative bands are evidence-strength judgements, not calibrated probabilities. Codenames stay opaque by design; the annex never discloses the underlying algorithms.
Machine Learning & Ensemble Clustering Interpretation
- Cluster Stability: high stability across varied initial conditions and resampling means higher confidence; low stability means treat the cluster as a hypothesis pending corroboration.
- Separation & Overlap: strong separation suggests meaningful operational differences; overlap may reflect shared providers or multi-tenant hosting. Avoid actor claims without supporting evidence.
- Ensemble Meaning: agreement across diverse model views strengthens trust; divergence indicates alternative plausible groupings and should drive hypothesis testing.
- Size & Distribution: large, diffuse clusters are often background archetypes; small but cohesive clusters are frequently specialized and operationally significant.
- Data Sparsity Signals: a missing feature can be deliberate hardening, a service that was never configured, or a collection gap, and the report uses cautious language accordingly. What separates an informative absence from a meaningless one is what the rest of the population does: a field most other hosts carry and this cohort does not is a distinction you can hunt on, and it appears as an exclusion clause in the cluster's query. A field almost nobody has anywhere distinguishes nothing and is dropped silently.
Outlier Handling Playbook (high vs low confidence)
-
High-confidence outlier (small but cohesive; consistent over time; strong unique features)
- Treat as a priority deep dive (specialized capability, test bed, or control plane).
- Actions: extract exemplars, craft narrow compound detections, add zero-match watcher queries.
- In the current format these often appear in 6.B (Method Vela) or as a noise cluster profile in 11.3: a tight, structurally-unique group separated from the main population is exactly the high-confidence case.
-
Low-confidence outlier (unstable labeling; inconsistent features; low assignment confidence)
- Treat as suspected noise or decoy until proven otherwise.
- Actions: recheck provider and ASN stratification, assess collection gaps and time skew, and run feature-family sensitivity tests (does it disappear when a feature family is withheld?). These map to the 6.A instability patterns.
Analyst Tips
- Treat "OK"-rated and low-stability clusters as hypotheses until corroborated by multiple fingerprints.
- Judge a cluster on its profile, not on its query's match count. A cluster with a thin query and a distinctive profile is a completely different finding from a cluster with a thin profile, and only one of them is a dead end.
- Prefer high-specificity fingerprints (certificate subjects, JA3, JA3S, JARM, tight product-stack combinations) for initial detections, and always deploy them as a compound, never a lone parameter.
- Run the reduced pivot first and the full conjunction second. Do not read a zero-match result on a hash-stacked conjunction as a uniqueness score; it is the expected outcome. Zero on a short, low-specificity query is the one that means something.
- Read the detection objective before the rule. "Expected match population" tells you what your alert queue will look like, and "defeated by" tells you when to stop trusting it.
- Take exclusion terms seriously. On a cohort defined by what it lacks, the negations are usually the terms doing the work, and adding one is often how you get a commodity pattern down to a workable population.
- Check the falsification condition before you commit time to an anchor. Knowing what would kill a lead is worth more than another reason to believe it.
- When a hypothesis is capped (non-discrimination), keep competing hypotheses active and drive targeted collection. Do not report a winner the data does not support.
- Always read 6.B. The asset that matters most is frequently the one the main clustering threw into noise.
- An IP flagged in both 6.A and 6.B is your single highest-priority investigation target.
- Outliers are double-edged: high-confidence outliers may reveal specialized or hidden capabilities (investigate first); low-confidence outliers can be noise or decoy nodes (investigate before discarding or acting).
- Verify that outlier traits are genuine (rare certificate subjects, distinctive provider use) rather than collection bias or boundary effects.
