OpAMP, the Open Agent Management Protocol, is how you stop SSH-ing into hosts to change an OpenTelemetry Collector config. With more than a handful of Collectors, the hard part is knowing what each one runs, whether itâs healthy, and getting a change such as âstop shipping that email attributeâ onto all of them without taking the pipeline down. This post shows how OpAMP does that for a Collector fleet, with a lab I ran locally and an OTTL redaction change pushed over the wire.
In February I went to the SRE NL meetup at Elastic in Amsterdam. Near the end of the OpenTelemetry talk, a âTo infinity and beyondâ slide listed Baggage, OTTL and OpAMP as the next things to learn, and OpAMP is described there as âremote management of large fleets of data collection Agentsâ. I wrote the evening up in my SRE NL at Elastic recap.

In the audience for the OpenTelemetry talk at SRE NL at Elastic Amsterdam.
Versions I verified against: OpAMP specification (Status: Beta, opamp-spec v0.20.0), opamp-go v0.25.0, OpenTelemetry Collector Contrib v0.162.0 and the OpAMP Supervisor v0.162.0. I ran everything below on macOS (arm64) with the release binaries.
What OpAMP is, and what it isnât
OpAMP is a network protocol. The specification says it lets agents report their status to a server, receive configuration from it, and receive agent package updates. Itâs vendor-agnostic, so one server can manage a mixed fleet.
The pieces:
- Server: holds the desired state and the fleet inventory. Many clients connect to one server.
- Client: speaks OpAMP on behalf of an agent. The spec doesnât require the client to live inside the agent. It can be a plugin, a sidecar or a separate Supervisor process.
- Transport: WebSocket or plain HTTP (polling). Servers should accept both. The default path is
/v1/opampand the default port is 4320. - Identity: each agent has an
instance_uid. The extensionâs docs ask for a UUIDv7.
The specification as a whole is marked Beta. Some fields, such as available components, custom messages and heartbeats, are still marked Development.

The slide at SRE NL at Elastic: OpAMP as âremote management of large fleets of data collection Agentsâ, next to Baggage and OTTL.
What it isnât: a rollout engine. OpAMP carries the config, the status and the hashes. Deciding which Collectors get which config, and when, is your serverâs job.
The capabilities that matter
During the first message exchange, the agent and server each send a capabilities bit field. A server must not use a capability the agent didnât announce. These are the ones that matter for running a Collector fleet:
| Capability | What it gives you |
|---|---|
ReportsStatus | Mandatory. Agent description: service name, version, OS, your own attributes. |
AcceptsRemoteConfig | The server can push a config. Status comes back as APPLYING, APPLIED or FAILED with the config hash. |
ReportsEffectiveConfig | The agent reports the config itâs actually running, after merging local and remote. |
ReportsHealth | Healthy or not, plus a status string. |
ReportsOwnMetrics / Logs / Traces | The server tells the agent where to send its own telemetry. |
AcceptsRestartCommand | The server can ask for a restart. |
AcceptsPackages / ReportsPackageStatuses | Package (binary) offers and install status. Beta in the spec. |
OpAMP extension vs OpAMP Supervisor
Collector Contrib ships two components, and the difference is the most important thing to understand.
The OpAMP extension (opamp, in extension/opampextension) is alpha and ships in the contrib and k8s distributions. It runs inside the Collector and reports: agent description, effective config, health and available components. It can also accept a restart command behind the extension.opampextension.RemoteRestarts feature gate. It doesnât accept remote config. When I connected a Collector with only the extension, it announced ReportsStatus|ReportsEffectiveConfig|ReportsHealth|ReportsAvailableComponents and nothing else.
The OpAMP Supervisor (cmd/opampsupervisor) is a separate binary, also alpha, published under release tags that start with cmd/opampsupervisor. It starts the Collector as a child process, connects to your OpAMP server, receives remote config, merges it with local files, writes the result and restarts the Collector. It configures the extension to connect back to a local OpAMP endpoint inside the Supervisor. Its README lists AcceptsRemoteConfig, effective config, own telemetry, connection settings, restart and available components as implemented. ReportsHealth works with caveats. Package handling and Collector binary updates are still linked to open issues.
My take: use the extension alone if you want an inventory and nothing pushed into production. Use the Supervisor if the server is meant to change configs.
A local lab: example server, Supervisor and a Collector
The opamp-go repository includes an example server with a small web UI. Itâs a reference implementation, not a product.
git clone --depth 1 --branch v0.25.0 https://github.com/open-telemetry/opamp-go.git
cd opamp-go/internal/examples/server
go run .
# OpAMP on 0.0.0.0:4320/v1/opamp with TLS, UI on http://localhost:4321The server uses demo certificates from internal/examples/certs, and its listener binds to every interface. Keep it on a laptop.
Get the Collector and the Supervisor (swap in your OS and architecture):
V=0.162.0
curl -fLO "https://github.com/open-telemetry/opentelemetry-collector-releases/releases/download/v${V}/otelcol-contrib_${V}_darwin_arm64.tar.gz"
tar xzf "otelcol-contrib_${V}_darwin_arm64.tar.gz" otelcol-contrib
curl -fL -o opampsupervisor "https://github.com/open-telemetry/opentelemetry-collector-releases/releases/download/cmd%2Fopampsupervisor%2Fv${V}/opampsupervisor_${V}_darwin_arm64"
chmod +x opampsupervisorA local base config, base.yaml, that the Collector always gets:
receivers:
otlp:
protocols:
http:
endpoint: 127.0.0.1:4318
exporters:
debug:
verbosity: detailed
service:
pipelines:
logs:
receivers: [otlp]
exporters: [debug]And supervisor.yaml:
server:
endpoint: wss://127.0.0.1:4320/v1/opamp
tls:
ca_file: ./opamp-go/internal/examples/certs/certs/ca.cert.pem
capabilities:
accepts_remote_config: true # default false: remote config is opt-in
reports_effective_config: true
reports_health: true
reports_own_metrics: false # default true; off here, nothing to receive it
agent:
executable: ./otelcol-contrib
config_files:
- ./base.yaml
validate_config: true # run `otelcol validate` before applying (default false)
automatic_config_rollback: true # fall back to the last working config (default false)
passthrough_logs: true # Collector logs in the Supervisor log
description:
non_identifying_attributes:
fleet.ring: canary
storage:
directory: ./supervisor-state # default /var/lib/otelcol/supervisor
healthcheck:
endpoint: 127.0.0.1:13140 # GET /health; disabled unless setRun ./opampsupervisor --config=supervisor.yaml. The Supervisor logged Connected to the OpAMP server., started the Collector, and the server saw the agent with AcceptsRemoteConfig in its capabilities, fleet.ring: canary in its description and RemoteConfigStatus=APPLIED. The agent then appears in the UI on port 4321.
Pushing an OTTL redaction change
Before any change, I sent an OTLP/HTTP log record with an email in the body and the attributes user.email, http.request.header.authorization and db.password. The debug exporter printed all of it, plus a DEBUG record. Thatâs the leak we want to fix fleet-wide.
The remote config only contains what changes. The Supervisor merges it on top of base.yaml:
processors:
filter/drop-debug:
error_mode: ignore
log_conditions:
- log.severity_number < SEVERITY_NUMBER_INFO
transform/redact-pii:
error_mode: ignore
log_statements:
- delete_key(log.attributes, "http.request.header.authorization")
- delete_matching_keys(log.attributes, "(?i).*(password|secret|token).*")
- set(log.attributes["user.email"], SHA256(log.attributes["user.email"])) where log.attributes["user.email"] != nil
- replace_pattern(log.body, "[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\\.[A-Za-z]{2,}", "[email]")
- replace_pattern(log.body, "\\b(?:\\d[ -]?){13,16}\\b", "[card]")
service:
pipelines:
logs:
processors: [filter/drop-debug, transform/redact-pii]Line by line:
filterdrops what matches. Any log below INFO is removed before the transform runs. The filter processorâs logs signal is alpha in v0.162.0.delete_keyanddelete_matching_keysremove the auth header and anything whose key looks like a credential. The(?i)makes the regex case-insensitive.SHA256replaces the email with its hash, so you can still count unique users without storing the address.replace_patternonlog.bodymasks emails and card-like numbers inside free text. Attribute rules donât help when the PII is in the message string.error_mode: ignorelogs a failed statement and moves on. The transform processorâs README recommends it, so one bad record doesnât drop the batch.
Before pushing, validate the merged result exactly as the Supervisor will build it:
./otelcol-contrib validate --config=base.yaml --config=remote-redact.yamlThen paste it into the agentâs page in the example UI and save. The server logged APPLYING, then APPLIED. Sending the same records again gave:
Body: Str(payment ok for [email] card [card])
-> user.email: Str(86e0b9e56c17cc4d12387e1949b85053fbe73bc3ce5a1188713a9d300cc6133d)
-> http.route: Str(/pay)The authorization header and password were gone, and the DEBUG record never reached the exporter.
Contrib also has a dedicated redaction processor; the OTTL version is easier to roll out as a small diff.
Rolling out config changes safely across a fleet
This is where fleets break. A safe rollout uses settings that are already in the Supervisor:
- Validate on the agent. With
validate_config: true, the Supervisor runs the Collectorâsvalidateagainst the merged config before switching. I pushed a config that called a non-existentdelete_keysfunction. The Supervisor loggedNew configuration failed validation, keeping previous config, reportedFAILEDwith the validation error to the server, and redaction kept working on the old config. - Roll back on a bad start.
automatic_config_rollback: truecaches the last working remote config and restores it if a new one stops the Collector from starting. Both settings are off by default. Turn them on. - Stage by ring. Put
fleet.ring(or region, cluster, team) inagent::description::non_identifying_attributes, and have your server offer a new config to canaries first. Promote only when they reportAPPLIED, healthy, and their effective config hash matches what you sent. - Survive an outage.
startup_fallback_configsgives a Collector a known-good, standalone config when it starts with no persisted state and canât reach the server. - Watch the Collectors, not just the server. With
reports_own_metrics, the server can point each Collectorâs own metrics at an OTLP endpoint. For an OTTL change, compareotelcol_processor_incoming_itemsandotelcol_processor_outgoing_itemsfor the filter before and after.

The OpenTelemetry talkâs live demo in Elastic Observability, on both screens at SRE NL.
My take: a config that passes validate can still drop everything. The canary ring and a throughput comparison catch that; the validator doesnât.
Security: TLS, auth and what remote config may do
The spec is blunt: remote configuration and downloadable packages are âa significant security riskâ, because a compromised server can make every agent do undesirable work. Treat your OpAMP server like your CI system.
- TLS everywhere. Both components use the Collectorâs standard client TLS settings. Use
ca_file(andcert_file/key_filefor mTLS) instead ofinsecure_skip_verify, which only belongs in the READMEâs local example. - Authenticate the client. The spec allows HTTP Basic or Bearer auth during the connection upgrade, and a server must answer 401 on failure. The extension takes
headersor anauthextension ID per transport. For example,auth: bearertokenauthwith abearertokenauthextension (beta) reading a tokenfilename. The Supervisor hasserver::headers. Itsserver::authneeds the alphaopampsupervisor.Extensionsfeature gate. - Keep remote config opt-in. The spec recommends it, and the Supervisor defaults
accepts_remote_configto false. - Pin what remote config canât override. Files in
agent::config_filesmerge top to bottom, and$REMOTE_CONFIGgoes last by default. The Supervisorâs design doc shows listing acompliance_config.yamlafter$REMOTE_CONFIG, so your local redaction pipeline wins. That matters because the README lists âSanitization or restriction of Collector configâ as not implemented yet. - Mind what you report.
reports_raw_configon the extension can expose secrets written in config files, and itâs off by default. The Supervisorâscollector_crash_log_snippet_kibis off by default too, since Collector logs may contain sensitive data. - Least privilege. The spec recommends not running the agent as root.
Pitfalls
- Expecting the extension to apply configs. It doesnât. Remote config needs the Supervisor or your own OpAMP client.
$in pushed OTTL. The merged config goes through the Collectorâs normal variable expansion, so a pushed"a$$b"arrived asa$b. Escape$as$$, as the OTTL docs say for any Collector config.- Deprecated
reports_remote_config. In v0.162.0 the Supervisor logs an error if you set it. Remote config status is reported wheneveraccepts_remote_configis on. - Lists replace, maps merge. My remote
processorslist was added to the local pipeline because the local one had none. Without the experimentalconfmap.enableMergeAppendOptiongate, a list in a later file discards the earlier list instead of appending to it. - False validation failures. The Supervisor warns that validation âmay fail for valid configs if resources (e.g., ports) are temporarily unavailableâ.
- Restart over SIGHUP. The extensionâs restart command uses
SIGHUP, which isnât supported on Windows. A bad config then leaves the Collector down until something else restarts it. - Hashing isnât anonymising. An unsalted SHA-256 of an email can be reversed with a list of known addresses. Delete it if you donât need it.
If the traces themselves are the problem, start with my OpenTelemetry trace quality checklist. Its gateway config is a good candidate to manage through OpAMP once itâs stable.

