Scanning, and the two switches that arm it
Active scanning is the only part of bitscanner that sends unsolicited packets to devices somebody else owns. Probing a network you are not authorised to probe is a legal question before it is a technical one — and the segment in front of the agent is chosen by the host's interfaces, not by you: a VPN, a bridged container network or a wide corporate interface can put address space in reach that nobody intended.
So the envelope, not the capability, is the default. Every switch below fails closed, and the ones that matter can only be set by a human editing a file.
For what the agent reports without any of this, see the bitscanner overview.
The two switches
scanning.enabledarms the envelope — "this agent may probe, here is the scope, and I am authorised for it".scanning.probes.<name>chooses which probes run inside it — how loud that is.
Every probe defaults to false, so scanning.enabled: true on its own still sends exactly
zero packets.
They are two decisions because the two failure modes are different: a wrong scope probes the wrong people, a wrong probe set probes the right people too aggressively. Widening either one is always an explicit edit.
scanning:
enabled: false # the envelope
i_accept_scanning_authorization: false # your assertion that you may probe these networks
allowed_cidrs: [] # the ONLY address space that may ever be probed
probes:
neighbor_discovery: false # …and every probe, individually
A probe armed while the envelope is off is a startup failure that names both keys, not a setting that is quietly ignored. Silently ignoring it would leave an operator reading their own file believing a scan is running — or concluding the product is broken when no results appear:
$ bitscanner config validate --config config.yaml
Error: configuration validation failed: failed to load config: config validation failed:
scanning.probes.neighbor_discovery, scanning.probes.port_scan are true but
scanning.enabled is false: the probe toggles choose WHICH probes run, scanning.enabled
arms the envelope that permits any of them at all, and with the envelope off every one of
those probes would be denied. Either set scanning.enabled: true (together with
scanning.i_accept_scanning_authorization and scanning.allowed_cidrs), or turn off
scanning.probes.neighbor_discovery, scanning.probes.port_scan
The same treatment applies to a probe whose prerequisite is off. A port scan runs over the neighbour list that neighbour discovery produces, so asking for one without the other is a configuration that would run and produce nothing — indistinguishable, from the operator's side, from a broken agent:
scanning.probes.port_scan is true but scanning.probes.neighbor_discovery is false: the
port scan probes the neighbours that neighbour discovery finds; with neighbour discovery
off the target list is empty and nothing would be scanned. Set
scanning.probes.neighbor_discovery: true as well, or turn off scanning.probes.port_scan
And enabling the envelope without the authorisation acknowledgement:
scanning.enabled is true but scanning.i_accept_scanning_authorization is false: active
scanning sends unsolicited packets to other people's devices, and setting this key is you
asserting that you are authorised to probe every network listed in scanning.allowed_cidrs
(you own them, or you hold written permission from whoever does). Probing a segment you
are not authorised to probe — a shared/colocated network, or a partner network reachable
over a VPN — may be unlawful.
There is no implicit scan-everything default either. allowed_cidrs must be non-empty when
the envelope is on, entries must be canonical networks, and a 0.0.0.0/0 is refused
outright:
scanning.allowed_cidrs entry "10.20.1.3/16" is not a network address: it authorises
10.20.0.0/16 (65536 addresses), which is almost certainly wider than intended. Write the
network explicitly, or use a /32 (or /128) for a single host
That refusal exists because every CIDR parser on earth silently widens a host address with a network mask. The operator wrote one address and authorised 65,534.
scanning.* and security.* can only come from the file
$ BITSCANNER_SCANNING_ENABLED=true bitscanner config validate --config config.yaml
Error: configuration validation failed: failed to load config: refusing to start:
BITSCANNER_SCANNING_ENABLED set in the environment, and the scanning envelope and the
security block are settable ONLY from the config file. Those keys — scanning.enabled,
scanning.i_accept_scanning_authorization, scanning.allowed_cidrs, every
scanning.probes.*, and security.i_accept_plaintext_egress — are acknowledgements that a
person is authorised to probe somebody else's network, or that a customer's network map
may cross the wire in cleartext. An acknowledgement has to be written by hand into a file
that can be reviewed and diffed, not inherited from a shell profile, a systemd drop-in, a
compose file or a parent process. This is NOT ignored and NOT applied: the agent stops.
This is the point of the acknowledgements, and it is worth being blunt about why.
The whole value of i_accept_scanning_authorization is that a person wrote it into a file
that can be reviewed, diffed and kept in configuration management. A variable exported by
a shell profile, a systemd drop-in, a compose file or a compromised parent process is none of
those things. If the environment could set it, the control would be a formality.
Three properties make that structural rather than a rule somebody has to remember:
- The environment is not a source for those keys. Environment binding is an explicit
allowlist of ordinary operational keys (
agent.id, intervals, batch sizes). Nothing in it can arm a probe or relax transport security. - The safety subtrees are re-read from the file by a second loader that has no environment binding of any kind, and those values overwrite whatever the main loader produced.
- An attempt is a refusal, not a filter. The whole
BITSCANNER_SCANNING_*andBITSCANNER_SECURITY_*namespace is reserved, and the error names the variable — because an operator who believes they armed a scan must not be told nothing.
The reserved namespace is deliberately wider than the config keys: a variable of your own
called BITSCANNER_SECURITY_TOKEN is refused too. A refusal that names the variable is
recoverable in seconds; a pattern list that has to guess which spellings matter is not.
The probe ladder
All seven default to false. Listed least intrusive first — and each line says what actually
goes on the wire, because an operator cannot consent to something described only by its name.
| Probe | What it puts on the wire | Needs |
|---|---|---|
neighbor_discovery | Nothing. Reads the kernel's own ARP/NDP cache and enriches each entry from a bundled offline MAC/OUI vendor table | — |
gateway_discovery | Nothing new. Routing table plus MAC attribution from the ARP cache and the offline OUI table. ⚠️ Not true in 0.1.3-ga and earlier — it also sent an ICMP echo to every gateway it found; see the warning below | — |
reverse_dns | One PTR query per unnamed neighbour. This is a packet | neighbor_discovery |
service_discovery | An mDNS query to 224.0.0.251:5353 and an SSDP M-SEARCH to 239.255.255.250:1900. One packet each — and every device on the broadcast domain sees them | neighbor_discovery |
gateway_probe | TCP connects to the gateway's own ports (80, 443, 22, 23, 53; then 8080/8443 for service characterisation, with banner reads), plus an ICMP echo | gateway_discovery |
active_arp_sweep | INTRUSIVE. Walks the usable range of every attached subnet and TCP-connects to a handful of ports on every address — hundreds to thousands of connection attempts per cycle | neighbor_discovery |
port_scan | MOST INTRUSIVE. ~19 common TCP ports on every discovered neighbour, plus banner reads on those that present one (21/22/25/110/143) | neighbor_discovery |
Three of those deserve more than a table row.
reverse_dns is a toggle because it was not one
Neighbour discovery is sold as the safe first run: read the kernel's tables, send nothing.
That claim used to be false. Every unnamed neighbour was reverse-resolved unconditionally,
and there was no configuration that turned it off — a scan-envelope observation with only
neighbor_discovery armed recorded two granted probes per cycle, which were these lookups.
A reverse lookup is a packet, and it is a more revealing one than it looks. It goes to whatever resolver this host points at — on a corporate segment, often one the customer's own security team runs and logs — and the query discloses precisely which of their hosts this agent has enumerated, one PTR at a time. On a segment where the agent is being trialled without the network team's knowledge, that is the disclosure, not the port scan.
It is off by default so the zero-packet mode is reachable by configuration.
gateway_probe needs privilege it may not have
The reachability check dials a raw ICMP socket (ip4:icmp), which requires root or
CAP_NET_RAW. Without it, the dial fails and the agent falls back to a UDP datagram —
and the fallback is gated by the guard separately, because an unauthorised target must not
become authorised just because the first method failed.
The gateway is somebody's router, and often the single device whose failure takes the whole
segment down. Being in your routing table is not authorisation to probe it, which is why
gateway_discovery (passive, answers "which router am I behind, and did its MAC change")
and gateway_probe (active) are separate keys.
gateway_discovery opened this raw socket too, in 0.1.3-ga and earlierThe row above says gateway_discovery puts nothing new on the wire, and the paragraph above
says the raw ICMP socket belongs to gateway_probe. In every release up to and including
0.1.3-ga, neither is true.
The reachability call sat in the branch taken when gateway_probe is off. So arming
gateway_discovery and deliberately declining gateway_probe — which is precisely how an
operator says "do not open a raw socket against my router, and do not make me grant
CAP_NET_RAW" — still opened one, once per gateway per cycle, with a UDP datagram to port
33434 behind it whenever the raw socket was refused.
What was never affected: that probe was still authorised by the scan guard, so it could only
ever reach an address inside your allowed_cidrs, and it was refused outright with
scanning.enabled: false. The defect is that a switch documented as off still acted — not that
anything escaped the envelope.
If you are running 0.1.3-ga or earlier with gateway_discovery armed, either accept that
your gateway is being pinged every cycle, or turn gateway_discovery off until you can upgrade.
In source the passive path now sends nothing at all: it reports reachability only from the
kernel's own resolved ARP entry, reports no round-trip time it did not measure, and the build
fails if any function other than the gateway-probe path can reach that socket — pinning the
socket to a function rather than to its callers is what let this survive two review passes.
active_arp_sweep is bounded by the file, not by the host
The subnets come from this host's own interfaces, so a VPN or a wide corporate interface
— not your configuration — decides what is in front of the sweep. allowed_cidrs and
min_prefix_length are what hold it in. min_prefix_length defaults to 22 (~1,022 hosts)
and is refused below /16 entirely: not because of packets, which the host budget bounds,
but because the sweep loop at /8 would walk 16 million addresses asking permission for
each one.
The measured zero-packet result
The claim "neighbour discovery on its own sends nothing" is the kind that has to be measured
rather than asserted. Run on an ordinary Linux host with scanning.enabled: true,
neighbor_discovery armed and every other probe off — the allowed range below replaced with
a documentation range, and nothing else altered:
WARN scanguard active scanning is ENABLED: this agent will send unsolicited probes
{"allowed_cidrs": ["10.20.0.0/24"], "allow_public_targets": false,
"min_prefix_length": 22, "max_hosts_per_cycle": 1024, "max_concurrency": 16,
"per_probe_delay": "50ms", "max_cycle_duration": "5m0s"}
WARN agent ACTIVE SCANNING ARMED: this agent will send unsolicited probes
{"armed_probes": ["neighbor_discovery"], "confined_to_allowed_cidrs": ["10.20.0.0/24"], …}
INFO neighbor_scanner neighbor discovery complete
{"neighbors_found": 116, "subnets": 27, "duration": "71.116454ms"}
INFO network_collector scan envelope
{"probes_allowed_since_start": 0, "probes_denied_since_start": 0,
"denied_by_reason_since_start": {}, "hosts_probed_this_cycle": 0,
"host_budget_remaining_this_cycle": 1024}
INFO network_collector collection cycle complete
{"queued_for_delivery": ["network_state", "neighbor_table", "neighbor_discovery"]}
Four consecutive cycles, 116–119 real neighbours across 27 attached subnets, and
probes_allowed_since_start stayed at 0 in every one of them. The scan guard was never
consulted, because there was nothing to ask it about: the neighbours came out of the kernel's
tables and the vendor names out of a bundled table. Three telemetry records were still
queued per cycle.
That is the shape of a first deployment: switch the envelope on, arm neighbor_discovery
only, and find out what is on the segment before you decide whether anything louder is
justified.
probes_allowed_since_start and probes_denied_since_start are process-lifetime totals;
hosts_probed_this_cycle and host_budget_remaining_this_cycle reset each cycle. They used
to be printed together under a heading that claimed all four were per-cycle, and on a running
agent the allowed count climbed 686 → 1142 across cycles while the host count stayed at 16 —
which reads as an escalating scan and was nothing of the kind. The guard keeps no per-cycle
decision counters, so the totals are labelled as totals rather than a per-cycle number being
invented for the log line.
scanguard: the decision happens where the packet leaves
Every probe is authorised immediately before the connection is attempted — not when the target list is built.
A filtered target list is a convention, and conventions are what future call paths quietly route around. Somebody adds a code path that builds its own list, or reuses an address from a response, and the filter that was applied three functions earlier does not apply to it. Putting the check where the packet leaves means a new call path either asks the guard or does not compile against the collector's plumbing.
The guard fails closed at every level: the zero-configured guard denies everything; a nil guard handed to a collector is replaced by deny-all rather than bypassed; an address that does not parse — or parses ambiguously — is denied; budget, pacing and deadline exhaustion deny rather than degrade.
Every decision, allow or deny, is recorded with a machine-readable reason:
| Refused | Reason recorded |
|---|---|
| Active scanning is off | scanning_disabled |
| The target will not parse | unparseable_target |
| Loopback, link-local, multicast, unspecified, the network or broadcast address of the allowed range | reserved_address |
Anything outside allowed_cidrs | outside_allowed_cidrs |
Public address space without allow_public_targets | public_target_not_allowed |
max_hosts_per_cycle reached | host_budget_exhausted |
max_cycle_duration reached | cycle_duration_exceeded |
A subnet wider than min_prefix_length, or outside scope | subnet_wider_than_min_prefix, subnet_outside_allowed_cidrs |
| A multicast group that is not a known discovery group | not_a_discovery_multicast_group |
| The agent is shutting down, or the cycle's context was cancelled mid-probe | context_cancelled |
Grants carry a reason too — in_scope_private, in_scope_public_explicitly_allowed,
link_local_discovery_group, subnet_within_sweep_limits — so the audit trail says why a
probe was permitted and not merely that it was.
The refusals hold even when the allowlist would otherwise match. That is exercised directly
by the guard's own test suite, which passes on the shipped commit: loopback (v4 and v6),
link-local (v4 and v6), multicast including the SSDP group, the unspecified address, the
limited broadcast address, and the network and broadcast addresses of the allowed range are
all denied while inside allowed_cidrs; an in-scope private host is granted without the
public opt-in; and an IPv4-mapped IPv6 address cannot dodge the allowlist.
0 probes sent, 4094 denied outside_allowed_cidrs is not a malfunction. It is the operator-
visible proof that the allowlist is the thing deciding.
The ceilings are part of the envelope
| Key | Default | Bound | What it holds back |
|---|---|---|---|
min_prefix_length | 22 | 16–32 | The widest subnet a sweep may walk |
max_hosts_per_cycle | 1024 | 1–65536 | Distinct hosts probed per cycle, across all subnets |
max_concurrency | 16 | 1–256 | Probes in flight at once, enforced by the guard itself |
per_probe_delay | 50ms | 0–10s | Minimum gap between probe packets — per probe, so a 19-port scan of one host is paced 19 times |
max_cycle_duration | 5m | ≤ 1h | Wall-clock ceiling on probing within one cycle |
allow_public_targets | false | — | Probing outside RFC1918 / RFC6598 / RFC4193 |
These are not tuning suggestions. A max_concurrency of 10,000 with a zero delay across a
/22 is a denial-of-service against the customer's own segment, delivered by a binary they
installed on our recommendation — so the values are range-checked at load and refused
outside it.
Next steps
- What leaves the host — what the discovered records actually contain, how they travel, and the limits this release ships with.
- bitscanner overview — what a cycle produces, what it costs, and how to verify the download.
- bitenforcer's guards — the same design instinct applied to a different blast radius: measure it, name it, and refuse rather than guess.
- Asset Management — where discovered devices become assets with owners.
War diese Seite hilfreich?