Skip to main content

Investigate

Investigate is the page for checking one thing, or a page full of things, against the indicator corpus by hand. Paste a value, press Check, and read the verdict. It is the surface for the moment a ticket names a domain and somebody wants to know whether anyone has seen it before. The automated side of the same question, the fleet's own alert stream checked continuously, is what a sensor does and is described in How the sidecar works.

What can be pasted​

Six types of indicator are recognised: an IP address, either IPv4 or IPv6, with or without a port that is kept as context, a domain, a URL, and a file hash as MD5, SHA-1 or SHA-256, told apart by their length.

Input copied out of a report or a ticket is usually defanged, and Investigate understands the common conventions, so there is nothing to clean up first: hxxp:// and hxxps:// become the scheme they stand for, [.], (.) and [dot] become a dot, [@], [:] and [/] become the character they wrap, and example . com is read as a domain when nothing else fits. The page shows the value it actually looked up under the verdict, next to the pasted value, so a rewrite is never silent. The same normalisation is applied when a source publishes a defanged value into the corpus, which is why a pasted hxxp:// and a published one land on the same record.

Two modes share the page, chosen at the left edge of the search bar. One indicator takes a single value. Paste many takes log lines, an email header, a report or a plain list, splits it into candidate values and checks every one it can read, saying how many indicators it read from how many lines and how many it had to skip. Pasting several lines into the single field switches to Paste many with the text kept, and Load example fills the box with a few sample lines. Switching mode clears the previous result.

Before the first lookup the page shows how many indicators of each type the corpus holds, each one a link to that type on the Indicators page, two values to try, one known bad and one allowlisted, and, when the fleet has reported sightings, the values it saw most recently. Any of them runs a lookup in one click.

The verdict​

Every value comes back with one of three verdicts, and the third exists so the second cannot be misread:

VerdictMeaning
Known badThe value is in the corpus, published by at least one source
Not in corpusNo source has listed it. That is not the same as safe
AllowlistedShared or reserved infrastructure Pharos refuses to publish as an indicator, whatever a source says. A hosting platform's shared domain reads as "we will not call this an indicator" rather than "we checked and it is fine"

For a value that is known, the verdict carries the value with a Copy button, its type, its class, one of phishing, command and control, malware distribution, attack infrastructure or sinkhole, and the malware family and target brand where a source attributed them. Three cards follow:

  • Confidence, a score out of 100 on a gauge.
  • Corroboration, how many independent sources list the value and how many of Pharos's own sensors have seen it, zero included.
  • Lifetime, first seen, last seen and when the record expires unless it is seen again, on one line. A record past its expiry reads Expired. A reference link sits under it where the source published one.

Under the cards, Watch this family and Watch this brand open the watchlist builder already filled in, when the value has a family or a brand, and Browse similar opens the Indicators page on the same type and class.

note

A sinkhole is a domain seized by researchers or a CERT to neutralise a botnet. Do not block it: traffic reaching a sinkhole means one of the fleet's hosts is infected, and the sinkhole is what is stopping it from doing harm.

Which sources listed a value is not part of the record shown. The number of sources is, because it is what confidence rests on. The names are kept corpus-wide on the Sources page, which is also where the reason for that is.

When several values are checked at once they come back as a table, known first, then allowlisted, then not in corpus. Above it, one tile per verdict gives the count and its share, and clicking a tile shows only those rows. A fourth tile counts what could not be read. A known row opens in place onto the same cards as a single lookup. Copy known bad copies the known values, one per line, ready for a block list.

Nothing is stored, and nothing is billed​

Investigate is a read. What is pasted is matched against the corpus and discarded: it is never written to the workspace, never added to the corpus, and there is no history of past lookups on the page. The Clear button empties the form, and closing the page does the same.

It is also free, and this is a decision rather than an omission. A signal, the billed unit described in Signals, is a message from a sensor that Pharos checked. A person pasting a hash into the console is not a sensor. Investigate is therefore the surface that keeps working when a fleet has stopped at its monthly limit: the sensors go quiet, the analyst does not. A single lookup may carry up to a thousand values.

The indicator corpus​

The corpus that Investigate reads is built from public and openly licensed indicator feeds: government advisories, vendor research feeds, phishing and malware trackers, and CERT warning lists. The Sources page in the console names each one that is contributing today, each one approved and waiting on a collector, and carries the notices their licences ask for. Feeds are fetched once a day, and a source that fails on a given day is skipped rather than being allowed to stop the others.

Every indicator carries a lifetime measured from when it was last seen, so an address that no source has mentioned for a month drops out while a file hash stays for a year. A value that a source publishes again is refreshed rather than duplicated, and the count of independent sources is what raises its confidence.

Some values a source publishes are held back rather than counted. Reserved and special-purpose IP ranges, the shared domains of large cloud and content-delivery providers, and Pharos's own domain are quarantined: kept in the corpus but withheld from the published indicator totals, so they cannot inflate them, and never returned to a fleet as a match. Matching is by IP range for the address blocks and by domain suffix for the providers, so a lookup of one of these returns Allowlisted rather than Known bad, which is the same rule seen from the Investigate page.

To be told when the corpus gains an indicator that matches a subject worth watching, such as a malware family, an owned brand, or a whole type, create an indicator watchlist. It is described in Watchlists and advisories.

Browsing the indicators​

The Indicators page lists the corpus in five tabs. All indicators is every published indicator, Seen by sensors keeps the ones a Pharos sensor has observed, and Quarantine shows what the allowlist withheld. By family and By target brand list each malware family and each impersonated brand with its indicator count, its sensor hits and when it was last seen. Selecting a row there opens the list filtered to that family or brand.

The search box on the list is an exact lookup of one value, not a substring search, and a defanged value such as hxxp://example[.]com finds the same indicator as the plain one. The search box on the family and brand tabs matches any part of the name.

The Filters button narrows the list by:

  • type, class, platform and kind, each taking several values.
  • confidence: a minimum, a maximum or both, from 0 to 100.
  • sources: how many independent sources corroborate the indicator, as a minimum, a maximum or both.
  • expiring: indicators whose lifetime runs out within 7 or 30 days.
  • attribution: indicators that carry a family or a target brand, or neither.
  • seen since: indicators last seen on or after a given day.

Several values of one filter widen the results, so IP and Domain together match an indicator of either type. Different filters narrow them. Each filter in use shows as a chip under the search box, and its X removes it.

Clicking a column header sorts the whole result on that column, highest or newest first, and A to Z for the indicator, family and brand columns. A second click reverses the order and a third returns to the default, which is the most recently seen first on the list and the most indicators first on the family and brand tabs. The list sorts by indicator, confidence, sources, sensors, last seen and expiry, and the family and brand tabs by name, indicators, sensor hits and last seen. Indicators that never expire sort last on the first click on Expires, and first when the order is reversed. The Target brand and Expires columns are hidden by default and Columns shows them.

Every tab shows 20 rows per page. The tab, the search, the filters, the sort and the page are part of the page address, so a copied link opens the same view and the browser's Back button undoes the last change.