Zum Hauptinhalt springen
Version: Next 🚧

How bitmapper captures

This page describes the capture command's actual pipeline: what it opens, what it filters, what it remembers, and what each knob costs you.

Everything below happens on the host. The packets themselves never leave it — what leaves is a flow record summarising each tracked connection. See what leaves the host.

interfaces → BPF filter (kernel) → worker pool → connection tracker → flow record → platform
│ │
│ └→ one last
│ attribution attempt
├→ process attribution, retried on
│ later packets with backoff
└→ statistics (structured JSON log)

Choosing interfaces​

bitmapper interfaces lists what this host can capture on. Then either name them, or name none and let the agent choose:

capture:
interfaces: ["eth0", "eth1"] # empty = every interface that is up and not loopback

With capture.interfaces empty, the agent captures on every interface that is up and is not loopback. On a host with several NICs, a bridge, and container virtual interfaces, that is usually more than you meant — the same packet can be seen more than once as it crosses a bridge. Name the interfaces you actually care about.

capture.exclude_interfaces removes names from whichever list you end up with:

capture:
interfaces: [] # auto-detect: every interface up and non-loopback
exclude_interfaces: ["docker0", "veth0"]

It applies both to an explicit capture.interfaces list and to the auto-detected set, and matching is case-insensitive with surrounding whitespace ignored. That makes it the right tool for the common case: let the agent find the interfaces, then subtract the bridges and virtual interfaces you do not want doubled up.

Filtering: four keys, one kernel filter​

Four keys decide which packets bitmapper ever sees, and all four are compiled into a single BPF expression that is handed to libpcap before capture starts. Whatever they exclude is dropped by the kernel — not filtered out of the output afterwards, not copied into the agent at all.

KeyDefaultWhat it contributes
capture.filter"" — noneA raw BPF expression, used verbatim, wrapped in parentheses
capture.protocols["tcp", "udp"]An allow-list: (tcp or udp). Accepts only tcp, udp, icmp, all
capture.ports[] — noneAn allow-list: (port 80 or port 443)
capture.exclude_ports[] — noneA deny-list: not (port 22)

The clauses are joined with and, in that order. The configuration bitmapper ships (configs/bitmapper.yaml) sets protocols: [tcp, udp] and exclude_ports: [22], which compiles to:

(tcp or udp) and not (port 22)

Set all four and you get one expression — for example filter: "host 10.0.0.1", protocols: [tcp], ports: [443], exclude_ports: [22] compiles to (host 10.0.0.1) and (tcp) and (port 443) and not (port 22).

The agent logs the compiled expression at start-up whenever it differs from what you wrote in capture.filter, so you can read back exactly what the kernel was given rather than infer it. A port outside 1–65535 in either list is refused at start-up rather than failing later inside libpcap.

Correction: these three keys are live, and this page previously said they were not

An earlier version of this page carried a warning headed "Three filter keys that are read by nobody", which stated that capture.protocols, capture.ports and capture.exclude_ports were validated and then ignored. That was wrong. All three are compiled into the BPF filter described above and enforced by the kernel.

It was wrong in the direction that costs you, and capture.exclude_ports is the reason to re-read your configuration now. The shipped configuration contains exclude_ports: [22] — "do not capture SSH" — and that exclusion is real: SSH packets are discarded before they reach the agent. If you read the old warning and removed the key as dead weight, you switched off a working privacy control and began collecting traffic you had deliberately excluded. Nothing would have warned you, because from the agent's point of view you simply asked for more traffic.

The old warning's counter-example was also exactly backwards. capture.ports: [443] does not capture everything. With the default protocols it compiles to (tcp or udp) and (port 443) and captures only port 443:

# This captures TCP/UDP port 443 and nothing else — the ports key is applied, in the kernel
capture:
ports: [443]

The default is tcp + udp, so ICMP is not captured​

capture.protocols defaults to ["tcp", "udp"] even when your configuration file never mentions it. A host running a stock configuration therefore never sees ICMP, ICMPv6, SCTP, GRE, ESP or AH: the kernel drops them, the agent is never handed them, and no counter records what was filtered away. An absence and a silence look identical in the output.

That is a genuine blind spot if you are using bitmapper to see ping sweeps, tunnelled traffic or IPsec. To remove it, ask for no protocol constraint at all:

capture:
protocols: ["all"]

Two details that are easy to get wrong:

  • protocols: ["icmp"] covers both families — it compiles to (icmp or icmp6), so asking for ICMP does not quietly give you IPv4 only.
  • There is no value for SCTP, GRE, ESP or AH. protocols accepts tcp, udp, icmp and all and nothing else; anything else is refused at start-up with invalid protocol: … (valid: tcp, udp, icmp, all). If you need one of those specifically, set protocols: ["all"] and express the protocol in capture.filter.

capture.exclude_ports on its own does not create this blind spot. It renders as not (port 22), and a BPF port test is also true for packets that have no ports, so excluding SSH does not drop ICMP along with it. Only the protocol allow-list does that.

Prefer the narrow filter, whichever key expresses it​

BPF is worth learning either way. tcp and not port 22, port 443 or port 8443, not net 10.0.0.0/8 — each one is applied before the packet costs you anything, and capture.filter remains the only key that can express host, network or direction constraints. The three list keys are the readable shorthand for the protocol and port cases; they end up in the same place.

What is inert here is the logging flags​

--log-level and --log-format do nothing

This page used to point at the filter keys as the things that parse and have no effect. The keys that actually behave that way on the command line are the logging flags.

--log-level and --log-format are registered on the root command and appear in --help with defaults of info and json, but no code reads either one. The logger is a fixed zap production logger constructed at start-up — JSON, info level — and neither flag can change it. Passing --log-level debug gets you no extra detail; passing --log-format console still gets you JSON.

The other capture flags behave the same way. --interfaces, --filter, --snaplen, --promiscuous, --buffer-size, --workers, --queue-size and --stats-interval are bound into Viper under their flag names while the configuration is read from nested keys (capture.filter, capture.snaplen, performance.workers, …), so they are parsed and ignored. --enable-grpc and --enable-http are worse than ignored: they default to true in --help and describe servers this agent never constructs — there is no gRPC or HTTP listener to enable.

--config is the only flag that changes what the agent does. Put everything else in the YAML file.

Reading packets​

KeyDefaultWhat it does
capture.snaplen65535Bytes captured per packet (validated 0–65535)
capture.promiscuoustruePromiscuous mode
capture.buffer_size104857600Kernel capture buffer, in bytes
capture.timeout100msPacket read timeout

Snaplen is the lever most worth turning down. bitmapper builds a connection map — who talks to whom, over what, how much — and that lives entirely in the packet headers. Capturing the full 65535-byte payload copies application data into the agent's memory that nothing here reads. A snaplen in the low hundreds of bytes keeps every header the tracker needs and stops payload from being copied at all, which is both cheaper and a far smaller exposure if the host is compromised.

Promiscuous mode makes the NIC accept frames not addressed to it. On a switched network that mostly yields broadcast and multicast; on a mirror/SPAN port it is what makes the capture useful at all. It also requires elevated privileges and is worth turning off when you only want this host's own conversations.

Process attribution​

capture:
enable_process_mapping: true
process_mapping_interval: 5s
This section describes the NEXT release, not the binary you can download

Everything below — per-flow attribution, the backfill, the listening-socket inference and the attribution field — is unreleased. It is not in 0.1.0-ga (commit 64ebc64), which is still the only published binary. Run bitmapper version and compare before you plan around anything on this page.

What the binary you can download today does. Read from the code at commit 64ebc64:

  • Attribution is attempted once, when the connection is created, from a socket-table snapshot that can be up to process_mapping_interval old. There is no retry of any kind.
  • There is no listening-socket inference of any kind. The code that performs it does not exist at that commit. 0.1.0-ga can never emit listening_socket, and can never name a process it did not match on the flow's own socket.
  • The answer is always recorded on the source end. destination_process is never assigned anywhere in that release, so it can never appear in an exported record.
  • The exported process object carries pid, name, executable and user — and no attribution field.

An outbound connection's socket cannot be in a snapshot taken before that connection existed, so flows from 0.1.0-ga commonly carry no process at all. Measured on the development host: 0 of 12 flows attributed. Expect that from this release rather than reading it as a misconfiguration.

What the next release does. The attempt is retried on the flow's later packets and once more on the export path; both ends are attributed; and every process object carries an attribution field saying whether the identity came from the flow's own socket or from a listening socket. That inference is gated three ways — the endpoint must be an address on this host, exactly one process must be proved to hold the listening socket, and the flow must be TCP.

This is what separates a capture from a packet dump: a tracked connection is resolved to the PID, process name, executable path and user that owns the socket, by reading the host's socket tables.

Attribution is per flow, not per packet, and it is recorded against the end of the flow that owns the socket — so a flow record can carry source_process, destination_process, or neither. On a host that is one end of the conversation, the other end is a remote peer whose process is not on this machine and never will be; only one end is expected to be filled.

It is backfilled, not resolved once​

A flow's first attribution attempt happens when the connection is created, and it very often fails — an outbound connection's socket cannot appear in a socket-table snapshot that was taken before the connection existed. So the attempt is repeated:

  • On the flow's later packets, on a backoff that starts at 100 ms and doubles to a ceiling of 30 s. A flow that genuinely has no local socket — traffic the host is merely forwarding — therefore settles at a couple of lookups a minute rather than one per packet.
  • On the export path, before a record for the flow is built. For a flow that has already finished this is its only remaining chance: a local request/response can live well under a millisecond, so its socket is gone microseconds after its one packet-path attempt.
process_mapping_interval is not the dial for unattributed flows

This page previously described a trade-off in which a short interval attributes short-lived connections more often and a long one loses them. That is no longer how attribution works, and it now contradicts the configuration the agent ships — configs/bitmapper.yaml states the opposite in a comment on the same key.

process_mapping_interval sets the scheduled socket-table refresh. It is no longer the only thing attribution depends on: a lookup that misses the snapshot triggers an on-demand re-read of the kernel socket tables for that one flow, and the attempt is retried as described above. Lowering the interval is not the fix for flows that arrive unattributed, and raising it does not cost you the short flows it used to.

Connections that cannot be attributed are still tracked and still exported; they simply carry no process identity. That is a correct answer, not a gap — see below for why it is preferable to the alternative.

How the attribution was made: socket vs listening_socket​

Every process object in an exported flow record carries an attribution field saying how that identity was determined. It is a confidence signal, and the two values are not interchangeable:

"source_process": {
"pid": 41207,
"name": "curl",
"executable": "/usr/bin/curl",
"user": "deploy",
"attribution": "socket"
}
ValueHow it was determinedHow far to trust it
socketMatched this flow's own socket in the host's socket tables, on the full four-tupleA fact. This process held that socket.
listening_socketNo socket for this flow was found. The identity was taken from the process listening on that address and port — and only where that address is on this host, exactly one process is proved to hold the socket, and the flow is TCPAn inference, and labelled as one. Guarded, but still a weaker statement than socket: see the sole-owner rule and TCP only.

The weaker value exists because the alternative is nothing at all: a flow that lives a few hundred microseconds is closed long before any read of /proc could see its socket, and the listening socket is the only thing that outlives it. It is worth reporting — but it must not be read as the same statement as socket.

A consumer that treats the two alike is drawing a conclusion the agent did not make. If you build alerting, dependency maps or policy on top of flow records, branch on this field.

Correction: this page described an over-applied inference that no release contains

An earlier version of this page carried a danger block headed "listening_socket is an inference, and today it is over-applied". It made two specific claims — that the inference "is not checked against the local host", and that it "does not identify which listener" — and told you to discard listening_socket outright for any endpoint that is not an address on the reporting host.

Both claims are false of every version of bitmapper you can obtain. They described an internal state that was corrected before anything was published:

  • The published 0.1.0-ga has nothing to over-apply. It contains no listening-socket inference at all, and cannot emit listening_socket.
  • The next release checks both of those things before it infers. The endpoint must be an address on this host — which is what stops an outbound flow to a remote peer being credited to a local wildcard listener — and exactly one process must be proved to hold that listening socket, which is what stops one of a pre-forking server's siblings being named. Measured after the fix, on the same host and the same traffic that produced the fabrications: no invented attributions at all, and every attribution the agent did make was correct.

The advice that followed from those claims was sound reasoning about code that no customer has. listening_socket is still the weaker of the two values and is still worth branching on — it is an inference, and it is labelled as one — but it is a guarded inference rather than an unchecked one, and "discard it for a non-local endpoint" now describes a rule the agent applies itself, before the record is built.

The inference is TCP only, so a UDP flow never carries it​

listening_socket can only ever appear on a TCP flow. UDP has no LISTEN state: a bound UDP socket is indistinguishable from a server's, so there is nothing to infer from. The inference returns nothing for any protocol other than TCP, before it examines anything else.

That is not a limitation waiting to be lifted — it is what prevents a specific wrong answer. A UDP flow can perfectly well land on a port that a TCP server listens on; 53 and 853 are the ordinary case. Without the protocol check, a DNS server's TCP listener would be named as the owner of unrelated UDP traffic on the same port.

So a UDP flow whose own socket could not be found is exported with no process, and that is the ordinary outcome for UDP rather than an unusual one: a single DNS query is over in well under a millisecond, its socket is gone before any read of /proc could reach it, and there is no weaker answer to fall back on. Read a blank process field on a UDP flow as the only honest answer available for that protocol, not as a defect.

Direct attribution is unaffected by the protocol check. UDP sockets are read into the same socket tables, so a UDP flow whose own socket is present there with a matching four-tuple is attributed exactly as a TCP flow is, with "attribution": "socket". The catch is that a UDP socket that was never connect()ed has no remote address to report, so there is no four-tuple to match — which is the other reason UDP attribution is sparser than TCP's.

An inference needs a socket with exactly one owner​

bitmapper takes an identity from a listening socket only when it can prove that exactly one process holds that socket. Where it cannot prove that, the flow is exported carrying no process rather than one of the candidates.

A listening socket with several owners is the ordinary case, not an exotic one. Measured on the development host: one nginx listening socket on 0.0.0.0:443 — a single socket inode — was held by 17 processes, which is what any server that pre-forks its workers looks like. One socket, seventeen names that could be written into the record. Inbound flows to a socket like that are left unattributed, deliberately. Naming one of the seventeen would put a guess in the same field, and in the same shape, as a measurement.

Only the full periodic /proc refresh can produce that proof. It walks every process, so it can establish that a socket inode appears in one process's descriptors and in no other's. The cheaper on-demand walk — the one a missed lookup triggers for a single flow — stops at the first process holding the inode it is chasing: it can prove that a process holds a socket, but never that it is the only one.

So a listener that has only just started is not inferred from until the next full refresh has covered it. For a short time after a service starts, short-lived traffic into it can appear with no process attribution at all. Two properties bound that, and both matter when you are deciding whether you are looking at a defect:

  • It costs attribution, never correctness. No wrong process is produced in that window; the field is empty rather than misleading.
  • A listening socket outlives the refresh interval comfortably, so this is a startup-window effect and not a steady-state one. Once one full refresh has covered a listener, it stays covered for as long as it is listening.

How long the window lasts depends on process_mapping_interval and on the host, so there is no single number to quote. The refresh loop is duty-limited — walking /proc must not itself become the load on a busy machine — so a host with a large process table takes longer to finish a refresh, and the window is correspondingly longer there.

None of this stops a flow being captured, tracked, counted or exported. If you are looking at records with an empty source_process and destination_process, read the periodic statistics the agent logs before you conclude the agent is broken: packet and connection counters that keep moving tell you capture is working, and that what you are seeing is attribution being withheld.

A wrong attribution is worse than none: an unattributed flow is visibly incomplete, while a flow naming the wrong process is a confident false statement that a consumer has no way to detect. That is why an unattributed flow is left unattributed rather than guessed at, and why the inference is labelled rather than silently blended in with the facts.

The worker pipeline​

KeyDefaultConstraint
performance.workers4≥ 1
performance.queue_size100000≥ 100

Packets are handed to a pool of workers through a bounded queue. When the queue is full, the packet is dropped and the drop counter is incremented — the agent does not block the capture path to keep up, because blocking there would make the kernel's own buffer overflow instead.

The drop counter is the number to watch. It appears in the periodic statistics, and a persistently rising drop count means the host is seeing more traffic than this configuration can process. In order of effect, the fixes are: tighten the kernel filter — any of capture.filter, capture.protocols, capture.ports or capture.exclude_ports, since all four land in the same expression — then lower capture.snaplen, then raise performance.workers.

Connection tracking​

KeyDefaultConstraint
performance.connection_table_size1000000≥ 1000 — max tracked connections before eviction
performance.connection_timeout5m≥ 1s — idle time before a connection is collected
performance.gc_interval1m≥ 1s — how often collection runs

The tracker matches both directions of a conversation into one connection, runs a TCP state machine (SYN_SENT → ESTABLISHED → … → CLOSED), and counts packets and bytes per connection.

Two bounds keep it finite, and they are different mechanisms:

  • connection_table_size is a hard cap. Past it, entries are evicted — a host under a connection flood loses old entries rather than growing until the process is killed.
  • connection_timeout + gc_interval collect connections that have simply gone idle, whether or not the table is near its cap.

A host that makes many short-lived outbound connections — a busy web crawler, a CI runner — will churn this table hard. That is the intended behaviour: a bounded table that forgets is better than an unbounded one that ends the process.

Statistics​

performance:
stats_interval: 10s # 0 disables the reporter entirely

Packet, byte, drop and connection counters are written as structured JSON through the agent's logger on this interval. They are the local evidence that capture itself is working, independently of whether export is reaching the platform:

bitmapper capture --config /etc/bitmapper/config.yaml

The output is JSON at info level and there is no flag that changes that — see the logging flags. Pipe it through jq if you want it readable.

Watch the drop counter across several intervals. Zero drops and a rising connection count means the pipeline is keeping up. If the packet counter stays at zero on a host you know is busy, check the compiled filter the agent logged at start-up before you suspect the capture path — the protocol default excludes more than most people expect.

What leaves the host, and what does not​

The traffic itself stays here. What leaves is a flow record — a summary of a connection, not the connection's contents. To be exact about it:

DataWhere it goes
Packet headers and payload up to snaplenAgent memory only, for the life of the packet's processing
Packet payload bytesNowhere. Never exported, never written to disk
Connection records, with process attributionThe in-memory connection table — and, as flow records, to the platform over https
How each attribution was made (socket / listening_socket)Carried in the flow record beside the process from the next release — see attribution confidence
The process command lineNowhere. Deliberately excluded: command lines routinely carry credentials as arguments
Packet/byte/drop/connection countersThe agent's own log output
Agent registration and heartbeatThe control plane, over https

A flow record carries the endpoints, ports, protocol, connection state, byte and packet counters, and — for whichever end of the flow owns the socket — the attributed process PID, name, executable path, user and — from the next release — the attribution marker saying how that identity was determined. That is the whole of it. If you need a subnet kept out of that, output.filter.exclude_cidrs drops the flow at intake — before a record is built — and it matches either endpoint. See the overview.

There is no storage.* implementation and no file export, so nothing is persisted by the agent. When the process stops, the connection table goes with it.

Next steps​

  • bitmapper overview — release, platforms, privileges, and the full list of keys that are validated but do nothing.
  • bitscanner — outward-facing discovery.
  • What is bits? — how the agents fit together.

War diese Seite hilfreich?