How bitmapper captures
This page describes the capture command's actual pipeline: what it opens, what it filters,
what it remembers, and what each knob costs you.
Everything below happens on the host. The packets themselves never leave it — what leaves is a flow record summarising each tracked connection. See what leaves the host.
interfaces → BPF filter (kernel) → worker pool → connection tracker → flow record → platform
│ │
│ └→ one last
│ attribution attempt
├→ process attribution, retried on
│ later packets with backoff
└→ statistics (structured JSON log)
Choosing interfaces
bitmapper interfaces lists what this host can capture on. Then either name them, or name none
and let the agent choose:
capture:
interfaces: ["eth0", "eth1"] # empty = every interface that is up and not loopback
With capture.interfaces empty, the agent captures on every interface that is up and is not
loopback. On a host with several NICs, a bridge, and container virtual interfaces, that is
usually more than you meant — the same packet can be seen more than once as it crosses a bridge.
Name the interfaces you actually care about.
capture.exclude_interfaces removes names from whichever list you end up with:
capture:
interfaces: [] # auto-detect: every interface up and non-loopback
exclude_interfaces: ["docker0", "veth0"]
It applies both to an explicit capture.interfaces list and to the auto-detected set, and
matching is case-insensitive with surrounding whitespace ignored. That makes it the right tool
for the common case: let the agent find the interfaces, then subtract the bridges and virtual
interfaces you do not want doubled up.
Filtering: four keys, one kernel filter
Four keys decide which packets bitmapper ever sees, and all four are compiled into a single BPF expression that is handed to libpcap before capture starts. Whatever they exclude is dropped by the kernel — not filtered out of the output afterwards, not copied into the agent at all.
| Key | Default | What it contributes |
|---|---|---|
capture.filter | "" — none | A raw BPF expression, used verbatim, wrapped in parentheses |
capture.protocols | ["tcp", "udp"] | An allow-list: (tcp or udp). Accepts only tcp, udp, icmp, all |
capture.ports | [] — none | An allow-list: (port 80 or port 443) |
capture.exclude_ports | [] — none | A deny-list: not (port 22) |
The clauses are joined with and, in that order. The configuration bitmapper ships
(configs/bitmapper.yaml) sets protocols: [tcp, udp] and exclude_ports: [22], which
compiles to:
(tcp or udp) and not (port 22)
Set all four and you get one expression — for example filter: "host 10.0.0.1",
protocols: [tcp], ports: [443], exclude_ports: [22] compiles to
(host 10.0.0.1) and (tcp) and (port 443) and not (port 22).
The agent logs the compiled expression at start-up whenever it differs from what you wrote in
capture.filter, so you can read back exactly what the kernel was given rather than infer it. A
port outside 1–65535 in either list is refused at start-up rather than failing later inside
libpcap.
An earlier version of this page carried a warning headed "Three filter keys that are read by
nobody", which stated that capture.protocols, capture.ports and capture.exclude_ports
were validated and then ignored. That was wrong. All three are compiled into the BPF filter
described above and enforced by the kernel.
It was wrong in the direction that costs you, and capture.exclude_ports is the reason to
re-read your configuration now. The shipped configuration contains exclude_ports: [22] — "do
not capture SSH" — and that exclusion is real: SSH packets are discarded before they reach the
agent. If you read the old warning and removed the key as dead weight, you switched off a
working privacy control and began collecting traffic you had deliberately excluded. Nothing
would have warned you, because from the agent's point of view you simply asked for more traffic.
The old warning's counter-example was also exactly backwards. capture.ports: [443] does not
capture everything. With the default protocols it compiles to (tcp or udp) and (port 443) and
captures only port 443:
# This captures TCP/UDP port 443 and nothing else — the ports key is applied, in the kernel
capture:
ports: [443]
The default is tcp + udp, so ICMP is not captured
capture.protocols defaults to ["tcp", "udp"] even when your configuration file never
mentions it. A host running a stock configuration therefore never sees ICMP, ICMPv6, SCTP, GRE,
ESP or AH: the kernel drops them, the agent is never handed them, and no counter records what was
filtered away. An absence and a silence look identical in the output.
That is a genuine blind spot if you are using bitmapper to see ping sweeps, tunnelled traffic or IPsec. To remove it, ask for no protocol constraint at all:
capture:
protocols: ["all"]
Two details that are easy to get wrong:
protocols: ["icmp"]covers both families — it compiles to(icmp or icmp6), so asking for ICMP does not quietly give you IPv4 only.- There is no value for SCTP, GRE, ESP or AH.
protocolsacceptstcp,udp,icmpandalland nothing else; anything else is refused at start-up withinvalid protocol: … (valid: tcp, udp, icmp, all). If you need one of those specifically, setprotocols: ["all"]and express the protocol incapture.filter.
capture.exclude_ports on its own does not create this blind spot. It renders as
not (port 22), and a BPF port test is also true for packets that have no ports, so excluding
SSH does not drop ICMP along with it. Only the protocol allow-list does that.
Prefer the narrow filter, whichever key expresses it
BPF is worth learning either way. tcp and not port 22, port 443 or port 8443,
not net 10.0.0.0/8 — each one is applied before the packet costs you anything, and
capture.filter remains the only key that can express host, network or direction constraints.
The three list keys are the readable shorthand for the protocol and port cases; they end up in
the same place.
What is inert here is the logging flags
--log-level and --log-format do nothingThis page used to point at the filter keys as the things that parse and have no effect. The keys that actually behave that way on the command line are the logging flags.
--log-level and --log-format are registered on the root command and appear in --help with
defaults of info and json, but no code reads either one. The logger is a fixed zap
production logger constructed at start-up — JSON, info level — and neither flag can change it.
Passing --log-level debug gets you no extra detail; passing --log-format console still gets
you JSON.
The other capture flags behave the same way. --interfaces, --filter, --snaplen,
--promiscuous, --buffer-size, --workers, --queue-size and --stats-interval are bound
into Viper under their flag names while the configuration is read from nested keys
(capture.filter, capture.snaplen, performance.workers, …), so they are parsed and ignored.
--enable-grpc and --enable-http are worse than ignored: they default to true in --help
and describe servers this agent
never constructs — there is no gRPC or HTTP listener to enable.
--config is the only flag that changes what the agent does. Put everything else in the
YAML file.
Reading packets
| Key | Default | What it does |
|---|---|---|
capture.snaplen | 65535 | Bytes captured per packet (validated 0–65535) |
capture.promiscuous | true | Promiscuous mode |
capture.buffer_size | 104857600 | Kernel capture buffer, in bytes |
capture.timeout | 100ms | Packet read timeout |
Snaplen is the lever most worth turning down. bitmapper builds a connection map — who talks to whom, over what, how much — and that lives entirely in the packet headers. Capturing the full 65535-byte payload copies application data into the agent's memory that nothing here reads. A snaplen in the low hundreds of bytes keeps every header the tracker needs and stops payload from being copied at all, which is both cheaper and a far smaller exposure if the host is compromised.
Promiscuous mode makes the NIC accept frames not addressed to it. On a switched network that mostly yields broadcast and multicast; on a mirror/SPAN port it is what makes the capture useful at all. It also requires elevated privileges and is worth turning off when you only want this host's own conversations.
Process attribution
capture:
enable_process_mapping: true
process_mapping_interval: 5s
Everything below — per-flow attribution, the backfill, the listening-socket inference and the
attribution field — is unreleased. It is not in 0.1.0-ga (commit 64ebc64), which is
still the only published binary. Run bitmapper version and compare before you plan around
anything on this page.
What the binary you can download today does. Read from the code at commit 64ebc64:
- Attribution is attempted once, when the connection is created, from a socket-table
snapshot that can be up to
process_mapping_intervalold. There is no retry of any kind. - There is no listening-socket inference of any kind. The code that performs it does not
exist at that commit.
0.1.0-gacan never emitlistening_socket, and can never name a process it did not match on the flow's own socket. - The answer is always recorded on the source end.
destination_processis never assigned anywhere in that release, so it can never appear in an exported record. - The exported process object carries
pid,name,executableanduser— and noattributionfield.
An outbound connection's socket cannot be in a snapshot taken before that connection existed, so
flows from 0.1.0-ga commonly carry no process at all. Measured on the development host:
0 of 12 flows attributed. Expect that from this release rather than reading it as a
misconfiguration.
What the next release does. The attempt is retried on the flow's later packets and once more
on the export path; both ends are attributed; and every process object carries an attribution
field saying whether the identity came from the flow's own socket or from a listening socket.
That inference is gated three ways — the endpoint must be an address on this host, exactly
one process must be proved to hold the listening socket, and the flow must be TCP.
This is what separates a capture from a packet dump: a tracked connection is resolved to the PID, process name, executable path and user that owns the socket, by reading the host's socket tables.
Attribution is per flow, not per packet, and it is recorded against the end of the flow
that owns the socket — so a flow record can carry source_process, destination_process, or
neither. On a host that is one end of the conversation, the other end is a remote peer whose
process is not on this machine and never will be; only one end is expected to be filled.
It is backfilled, not resolved once
A flow's first attribution attempt happens when the connection is created, and it very often fails — an outbound connection's socket cannot appear in a socket-table snapshot that was taken before the connection existed. So the attempt is repeated:
- On the flow's later packets, on a backoff that starts at 100 ms and doubles to a ceiling of 30 s. A flow that genuinely has no local socket — traffic the host is merely forwarding — therefore settles at a couple of lookups a minute rather than one per packet.
- On the export path, before a record for the flow is built. For a flow that has already finished this is its only remaining chance: a local request/response can live well under a millisecond, so its socket is gone microseconds after its one packet-path attempt.
process_mapping_interval is not the dial for unattributed flowsThis page previously described a trade-off in which a short interval attributes short-lived
connections more often and a long one loses them. That is no longer how attribution works,
and it now contradicts the configuration the agent ships — configs/bitmapper.yaml states the
opposite in a comment on the same key.
process_mapping_interval sets the scheduled socket-table refresh. It is no longer the only
thing attribution depends on: a lookup that misses the snapshot triggers an on-demand re-read of
the kernel socket tables for that one flow, and the attempt is retried as described above.
Lowering the interval is not the fix for flows that arrive unattributed, and raising it does not
cost you the short flows it used to.
Connections that cannot be attributed are still tracked and still exported; they simply carry no process identity. That is a correct answer, not a gap — see below for why it is preferable to the alternative.
How the attribution was made: socket vs listening_socket
Every process object in an exported flow record carries an attribution field saying how
that identity was determined. It is a confidence signal, and the two values are not
interchangeable:
"source_process": {
"pid": 41207,
"name": "curl",
"executable": "/usr/bin/curl",
"user": "deploy",
"attribution": "socket"
}
| Value | How it was determined | How far to trust it |
|---|---|---|
socket | Matched this flow's own socket in the host's socket tables, on the full four-tuple | A fact. This process held that socket. |
listening_socket | No socket for this flow was found. The identity was taken from the process listening on that address and port — and only where that address is on this host, exactly one process is proved to hold the socket, and the flow is TCP | An inference, and labelled as one. Guarded, but still a weaker statement than socket: see the sole-owner rule and TCP only. |
The weaker value exists because the alternative is nothing at all: a flow that lives a few
hundred microseconds is closed long before any read of /proc could see its socket, and the
listening socket is the only thing that outlives it. It is worth reporting — but it must not be
read as the same statement as socket.
A consumer that treats the two alike is drawing a conclusion the agent did not make. If you build alerting, dependency maps or policy on top of flow records, branch on this field.
An earlier version of this page carried a danger block headed "listening_socket is an
inference, and today it is over-applied". It made two specific claims — that the inference
"is not checked against the local host", and that it "does not identify which listener" —
and told you to discard listening_socket outright for any endpoint that is not an address on
the reporting host.
Both claims are false of every version of bitmapper you can obtain. They described an internal state that was corrected before anything was published:
- The published
0.1.0-gahas nothing to over-apply. It contains no listening-socket inference at all, and cannot emitlistening_socket. - The next release checks both of those things before it infers. The endpoint must be an address on this host — which is what stops an outbound flow to a remote peer being credited to a local wildcard listener — and exactly one process must be proved to hold that listening socket, which is what stops one of a pre-forking server's siblings being named. Measured after the fix, on the same host and the same traffic that produced the fabrications: no invented attributions at all, and every attribution the agent did make was correct.
The advice that followed from those claims was sound reasoning about code that no customer has.
listening_socket is still the weaker of the two values and is still worth branching on — it is
an inference, and it is labelled as one — but it is a guarded inference rather than an
unchecked one, and "discard it for a non-local endpoint" now describes a rule the agent applies
itself, before the record is built.
The inference is TCP only, so a UDP flow never carries it
listening_socket can only ever appear on a TCP flow. UDP has no LISTEN state: a bound UDP
socket is indistinguishable from a server's, so there is nothing to infer from. The inference
returns nothing for any protocol other than TCP, before it examines anything else.
That is not a limitation waiting to be lifted — it is what prevents a specific wrong answer. A UDP flow can perfectly well land on a port that a TCP server listens on; 53 and 853 are the ordinary case. Without the protocol check, a DNS server's TCP listener would be named as the owner of unrelated UDP traffic on the same port.
So a UDP flow whose own socket could not be found is exported with no process, and that is the
ordinary outcome for UDP rather than an unusual one: a single DNS query is over in well under a
millisecond, its socket is gone before any read of /proc could reach it, and there is no weaker
answer to fall back on. Read a blank process field on a UDP flow as the only honest answer
available for that protocol, not as a defect.
Direct attribution is unaffected by the protocol check. UDP sockets are read into the same
socket tables, so a UDP flow whose own socket is present there with a matching four-tuple is
attributed exactly as a TCP flow is, with "attribution": "socket". The catch is that a UDP
socket that was never connect()ed has no remote address to report, so there is no four-tuple
to match — which is the other reason UDP attribution is sparser than TCP's.
An inference needs a socket with exactly one owner
bitmapper takes an identity from a listening socket only when it can prove that exactly one process holds that socket. Where it cannot prove that, the flow is exported carrying no process rather than one of the candidates.
A listening socket with several owners is the ordinary case, not an exotic one. Measured on the
development host: one nginx listening socket on 0.0.0.0:443 — a single socket inode — was
held by 17 processes, which is what any server that pre-forks its workers looks like. One
socket, seventeen names that could be written into the record. Inbound flows to a socket like
that are left unattributed, deliberately. Naming one of the seventeen would put a guess in the
same field, and in the same shape, as a measurement.
Only the full periodic /proc refresh can produce that proof. It walks every process, so it
can establish that a socket inode appears in one process's descriptors and in no other's. The
cheaper on-demand walk — the one a missed lookup triggers for a single flow — stops at the first
process holding the inode it is chasing: it can prove that a process holds a socket, but never
that it is the only one.
So a listener that has only just started is not inferred from until the next full refresh has covered it. For a short time after a service starts, short-lived traffic into it can appear with no process attribution at all. Two properties bound that, and both matter when you are deciding whether you are looking at a defect:
- It costs attribution, never correctness. No wrong process is produced in that window; the field is empty rather than misleading.
- A listening socket outlives the refresh interval comfortably, so this is a startup-window effect and not a steady-state one. Once one full refresh has covered a listener, it stays covered for as long as it is listening.
How long the window lasts depends on process_mapping_interval and on the host, so there is no
single number to quote. The refresh loop is duty-limited — walking /proc must not itself
become the load on a busy machine — so a host with a large process table takes longer to finish
a refresh, and the window is correspondingly longer there.
None of this stops a flow being captured, tracked, counted or exported. If you are looking at
records with an empty source_process and destination_process, read the periodic statistics
the agent logs before you conclude the agent is broken: packet and connection counters that keep
moving tell you capture is working, and that what you are seeing is attribution being withheld.
A wrong attribution is worse than none: an unattributed flow is visibly incomplete, while a flow naming the wrong process is a confident false statement that a consumer has no way to detect. That is why an unattributed flow is left unattributed rather than guessed at, and why the inference is labelled rather than silently blended in with the facts.
The worker pipeline
| Key | Default | Constraint |
|---|---|---|
performance.workers | 4 | ≥ 1 |
performance.queue_size | 100000 | ≥ 100 |
Packets are handed to a pool of workers through a bounded queue. When the queue is full, the packet is dropped and the drop counter is incremented — the agent does not block the capture path to keep up, because blocking there would make the kernel's own buffer overflow instead.
The drop counter is the number to watch. It appears in the periodic statistics, and a
persistently rising drop count means the host is seeing more traffic than this configuration can
process. In order of effect, the fixes are: tighten the kernel filter — any of
capture.filter, capture.protocols, capture.ports or capture.exclude_ports, since all
four land in the same expression — then lower capture.snaplen, then raise
performance.workers.
Connection tracking
| Key | Default | Constraint |
|---|---|---|
performance.connection_table_size | 1000000 | ≥ 1000 — max tracked connections before eviction |
performance.connection_timeout | 5m | ≥ 1s — idle time before a connection is collected |
performance.gc_interval | 1m | ≥ 1s — how often collection runs |
The tracker matches both directions of a conversation into one connection, runs a TCP state
machine (SYN_SENT → ESTABLISHED → … → CLOSED), and counts packets and bytes per
connection.
Two bounds keep it finite, and they are different mechanisms:
connection_table_sizeis a hard cap. Past it, entries are evicted — a host under a connection flood loses old entries rather than growing until the process is killed.connection_timeout+gc_intervalcollect connections that have simply gone idle, whether or not the table is near its cap.
A host that makes many short-lived outbound connections — a busy web crawler, a CI runner — will churn this table hard. That is the intended behaviour: a bounded table that forgets is better than an unbounded one that ends the process.
Statistics
performance:
stats_interval: 10s # 0 disables the reporter entirely
Packet, byte, drop and connection counters are written as structured JSON through the agent's logger on this interval. They are the local evidence that capture itself is working, independently of whether export is reaching the platform:
bitmapper capture --config /etc/bitmapper/config.yaml
The output is JSON at info level and there is no flag that changes that — see
the logging flags. Pipe it through jq if you want it readable.
Watch the drop counter across several intervals. Zero drops and a rising connection count means the pipeline is keeping up. If the packet counter stays at zero on a host you know is busy, check the compiled filter the agent logged at start-up before you suspect the capture path — the protocol default excludes more than most people expect.
What leaves the host, and what does not
The traffic itself stays here. What leaves is a flow record — a summary of a connection, not the connection's contents. To be exact about it:
| Data | Where it goes |
|---|---|
Packet headers and payload up to snaplen | Agent memory only, for the life of the packet's processing |
| Packet payload bytes | Nowhere. Never exported, never written to disk |
| Connection records, with process attribution | The in-memory connection table — and, as flow records, to the platform over https |
How each attribution was made (socket / listening_socket) | Carried in the flow record beside the process from the next release — see attribution confidence |
| The process command line | Nowhere. Deliberately excluded: command lines routinely carry credentials as arguments |
| Packet/byte/drop/connection counters | The agent's own log output |
| Agent registration and heartbeat | The control plane, over https |
A flow record carries the endpoints, ports, protocol, connection state, byte and packet
counters, and — for whichever end of the flow owns the socket — the attributed process PID,
name, executable path, user and — from the next release — the attribution marker saying how
that identity was determined. That is the whole of it.
If you need a subnet kept out of that, output.filter.exclude_cidrs drops the flow at intake —
before a record is built — and it matches either endpoint. See
the overview.
There is no storage.* implementation and no file export, so nothing is persisted by the
agent. When the process stops, the connection table goes with it.
Next steps
- bitmapper overview — release, platforms, privileges, and the full list of keys that are validated but do nothing.
- bitscanner — outward-facing discovery.
- What is bits? — how the agents fit together.
Cette page vous a-t-elle été utile ?