What leaves the host
bitscanner ships a description of somebody's network off the machine it runs on. That makes two questions load-bearing: what is in it, and how does it travel. This page answers both, and then states the limits of this release plainly.
For what arms the probing in the first place, see Scanning, and the two switches that arm it.
What is in the telemetry
Everything bitscanner ships is about the network around the host, never the host's own contents.
| Record | Fields |
|---|---|
network_state | Routes only: destination, gateway, interface, metric, flags |
neighbor_table | Per neighbour: ip_address, hardware_addr, interface, state, type |
neighbor_discovery | Per neighbour: address, MAC, vendor (offline OUI lookup), device_type, is_reachable, first_seen/last_seen — plus hostname if reverse_dns is armed, services (ports, and banners if port_scan is armed), mdns_services/ssdp_info if service_discovery is armed. Plus the subnets scanned |
gateway_discovery | Per gateway: address, MAC, vendor, is_default, reachability, and — only with gateway_probe — the ports that answered and any banner they presented |
service_discovery | Responders: name, type, host, port, protocol, discovered_via, mDNS TXT records |
Two consequences worth being explicit about:
- The louder probes record other people's software versions. A banner from port 22 or 25
is a version string belonging to a device you may not own. That is not a side effect to
discover later; it is what
port_scanandgateway_probeare for, and it is why they are off by default and gated behind an authorisation acknowledgement. hostnameonly exists if you armedreverse_dns. Without it, a neighbour is an address, a MAC and a vendor guess.
What it deliberately does not collect
network_state used to carry interfaces (name, MTU, MAC, addresses), listeners (the
local socket inventory), an always-empty connections list and a dns_servers field that
nothing ever populated — a null field asserting a resolver inventory that was never
collected. All of them are gone, along with the code and types behind them.
The first three were bitcollector's network and port collectors restated. bitscanner does not collect host facts — no processes, no installed packages, no local accounts, no command lines. If a record about the machine itself is what you need, it comes from the collector, under the collector's privacy defaults.
Each record is wrapped in an envelope carrying a SHA-256 checksum over its header and
payload. That is an integrity check, not a signature — bitscanner has no equivalent of
bitcollector's signed, hash-chained evidence chain.
Egress is https, or the agent does not start
The rule is about the data, not the credential: a map of the customer's network crossing the wire in cleartext is readable and alterable by anyone on the path whether or not a token rides along.
telemetry.destinations[0]: endpoint http://es.example.com:9200 is not https — refusing to
start. Everything this agent collects about the customer's network would cross that link
in cleartext, readable and alterable by anyone on the path, credential or no credential.
Use an https endpoint; if this is a lab and you accept that the data is exposed, set
security.i_accept_plaintext_egress: true
security.i_accept_plaintext_egress is the single, file-only opt-out, it is announced in the
log at every start, and it is for labs. With a credential configured there is no opt-out at
all — a plaintext endpoint, or tls.skip_verify: true, is refused outright, because
anything that can answer for the endpoint collects the credential.
Three more refusals in the same family:
- A credential embedded in the URL is refused, with instructions to move it to the
authblock. Go turns userinfo into a liveAuthorization: Basicheader, it lands in every log line that prints the endpoint, and the agent's redaction cannot reach it. - Certificate material that would be silently discarded is refused. Setting
ca_cert/client_cert/client_keywhiletls.enabledis false is an error, because the stanza would read as mutual TLS while presenting nothing. headers:does not exist. The old sample config advertisedheaders: {Authorization: "Bearer YOUR-TOKEN"}and there was never a line of Go that read it — operators believed their telemetry was authenticated while it went out anonymously. A config containing that key is refused by name.
No redirect is ever followed
refusing to follow the redirect from <from> to <to>: this agent never follows redirects on
egress, because net/http would re-attach the credential when the hostname matches (even
downgrading https to http) and would replay the batch body on a 307/308 to a host the
server chose. Point the endpoint at its final URL instead
Both halves of that are properties of the Go standard library, not speculation. Its rule for
re-attaching Authorization compares the hostname only — not the scheme, not the port —
so a server answering 301 Location: http://same-host:80/ gets the bearer token back in
cleartext. A custom API-key header is not on the standard library's sensitive list at all, so
it is copied to any host over any scheme. And on a 307/308 the request body — the
customer's network inventory — is replayed to whatever host the redirect names.
The outbound client is not an *http.Client. Every field of that struct is exported, so
c.CheckRedirect = nil or a swapped transport would have removed both controls in one
assignment — and both edits were proven to leave the test suite completely green. The client
is wrapped in a type that exposes only Do and CloseIdleConnections, and the transport is
built from a TLS config rather than accepted from the caller, so there is no caller-supplied
round-tripper left to rewrite a request with.
The credential is applied at send time, not at request-construction time, and it fails closed: if the channel or the credential does not check out — unresolved, not https, certificate verification off — the batch is not sent. "Set the header if we happen to have one" is the shape that shipped unauthenticated telemetry once already.
Endpoints are redacted in logs
An endpoint is logged on every export failure, so redaction is deny-by-default rather than a
list of credential-looking names: every query value and the whole fragment are redacted,
with parameter names kept so the line is still diagnosable (api_key=[redacted] says which
parameter was set). Path segments that look credential-bearing are redacted too. If the URL
will not parse at all, the redactor returns a constant — never the caller's text, because a
URL that fails to parse is usually one whose password contains the characters that break the
parser.
That last point has a stated limit: the path heuristic can miss a short or word-like secret
in a path segment. The load-bearing control is that credentials belong in the auth block,
where they are typed, vetted and never formatted into a string.
Secret files
A token is only out of the config file if nobody else can read it. bitscanner config validate refuses to start on any of these, and the message names the offending path and the
fix. Verified by execution — the paths below are documentation paths substituted for the
ones in the real transcript:
# mode
control_plane: auth.token_file "/etc/bitscanner/control-plane.token" is mode 0644 —
readable by group or other; restrict it with `chmod 600 /etc/bitscanner/control-plane.token`
# any directory on the path, not just the one holding the file
control_plane: the directory "/opt/agent" on the path to auth.token_file
"/opt/agent/secrets/control-plane.token" is mode 0775 — group- or world-writable, so
another account can replace the file this agent reads. It is not the directory holding the
file — it is 1 level(s) above it — but another account can rename the whole subtree and
put its own in place. Restrict it with `chmod 755 /opt/agent` (or tighter), or move the
secret somewhere only root and this agent can write
The full set of refusals:
| Refused | Because |
|---|---|
| Mode readable by group or other | Every local account on the host can read the token |
| Owned by a third account | "Only the owner can read it" is worth nothing when the owner is somebody else |
| Not a regular file | A FIFO or device is not a secret file — and a plain open of a FIFO with no writer would wedge the agent at startup with no diagnostic |
| Reached through a symlink | The link's target lives in a directory the check would not have looked at |
Any group- or world-writable directory on the path — including /tmp | That account can rename the subtree and put its own file in place |
Both the resolved path's ancestors and the path as written are walked, and the
resolved path is proven by device and inode to be the very file whose descriptor is being
read. Each of those was a real hole, closed in a different wave: checking only the immediate
parent, missing a world-writable grandparent, and then the mirror image — walking only the
resolved chain, which accepted a 0600 file in a 0700 directory reached through a 0777
one, because the resolved chain is immaculate and the written chain is where the attacker
lives.
Finally, each secret comes from exactly one of: the inline value, a <field>_file, or a
<field>_env naming an environment variable. Two sources at once is an error rather than a
precedence rule — with precedence, an operator who adds token_file while leaving a stale
inline token in place cannot tell which one is on the wire.
Enrolment: the key never leaves the host
You mint a single-use enrolment token in the Cert-IX dashboard and put it in a file only the agent can read:
enrollment:
endpoints:
- "https://<your-cert-ix-agent-endpoint>"
enrollment_token_file: "/etc/bitscanner/enrollment.token" # chmod 600, owner-only
On the first start the agent:
- Generates an ECDSA P-256 key pair on the host. The private half is written
0600insideagent.data_dir, which is forced to0700— created and re-chmoded, becauseMkdirAllleaves an existing directory's mode alone and an upgrade into a0755data directory would otherwise keep it. The private key never leaves the machine. - Registers once, presenting the enrolment token in the
X-Agent-Enrollment-Tokenheader and its public key in the body. - Stores the agent JWT the gateway returns,
0600, alongside its expiry.
After that the token is spent: the file can be deleted, and a restart loads the stored
identity and does not register again. A second registration would be a 409 against a
consumed token, which would leave the agent dead. The JWT is renewed against the refresh
endpoint an hour before it expires, using the token itself — no enrolment token is involved.
It is presented only in X-Agent-Enrollment-Token, on exactly one request. Putting it in
auth.token_file is the misconfiguration that made onboarding impossible in an earlier
build: the ingest gateway wants an agent JWT, so every telemetry request 401s. Use
auth: {type: enrollment} on the destination instead — it presents the stored identity, and
there is nothing to paste.
The enrolment endpoint must be https under every configuration.
security.i_accept_plaintext_egress does not apply to it — the identity comes back over it.
Neither the private key nor the token is ever logged, formatted into an error, or put in a
log field: the identity type's own String() and GoString() print only the agent ID and
the expiry, so a stray %v cannot write this host's JWT into a log file for good.
Honest limits
Stated here rather than discovered later.
The Go standard-library advisories that affected 0.1.2 are fixed in 0.1.3
0.1.3-ga is built with go1.25.12, which carries the fix for both advisories that
0.1.2-ga (built with go1.26.4) was exposed to:
| Advisory | CVSS | Exploited in the wild? | What it is | Fixed in 0.1.3-ga |
|---|---|---|---|---|
CVE-2026-39822 (GO-2026-4970) | 7.8 High | No — not on CISA KEV, EPSS below threshold | Root escape via symlink plus trailing slash in os | ✅ |
CVE-2026-42505 (GO-2026-5856) | 5.3 Medium | No — not on CISA KEV, EPSS below threshold | Encrypted Client Hello privacy leak in crypto/tls | ✅ |
🪤 Worth knowing if you pin toolchains yourself: a higher Go version is not always the
fixed one. CVE-2026-39822 is fixed in go1.25.12 on the 1.25 line but not until go1.26.5
on the 1.26 line, so go1.26.4 — a numerically newer toolchain — still carried it. A single
"minimum version" floor cannot express a per-branch backport, which is why this release pins
an exact image and lets the scanner, not the version number, be the authority. Our own build
was refused once on precisely this before the pin was corrected.
If you are still running 0.1.2-ga, it carries the High advisory — upgrade. The
toolchain version is printed by bitscanner version and recorded in the release's CycloneDX
SBOM, so check the build in front of you rather than trusting this table forever.
"Delivered" means "a 2xx came back"
The agent reports what it queued and what was accepted:
INFO network_collector collection cycle complete
{"queued_for_delivery": ["network_state", "neighbor_table", "neighbor_discovery"]}
INFO telemetry telemetry batch accepted by destination {"status": 200, …}
A 2xx is the destination's word, not proof the record was stored, and the wording says
accepted rather than delivered on purpose. The agent cannot know more than that, and a
line that claimed otherwise would be inventing a guarantee. The queued_for_delivery line is
printed every cycle — including with scanning off — so that a delivery failure can never be
mistaken for a collector that is not running.
config validate cannot tell an enrolment token from a JWT
Both are opaque strings in a file, so a bitscanner config validate on a configuration whose
auth.token_file contains an enrolment token reports:
Configuration is valid.
…and every export then fails at runtime. Validation checks the file's permissions, its
ownership, every directory on the path to it, and that exactly one source is configured — it
does not, and cannot, check that the bytes inside are the right kind of credential. If your
telemetry is 401ing on a fresh install, this is the first thing to check.
Next steps
- Scanning, and the two switches that arm it — what each probe puts on the wire, and the guard that authorises it.
- bitscanner overview — what a cycle produces, and how to verify the download.
- bitcollector privacy — the equivalent defaults for host data: command lines off, no environment variables, no password hashes.
- Data subject rights — how Cert-IX handles requests about personal data.
¿Te resultó útil esta página?