Skip to main content

hermes-otel

OpenTelemetry for Hermes Agent

Fan LLM traces, tool calls, API requests, and token metrics out to any OTLP-compatible observability backend — Phoenix, Langfuse, LangSmith, SigNoz, Jaeger, Grafana Tempo and LGTM, Uptrace, OpenObserve, Parseable, Honeycomb, W&B Weave, Elastic, OpenLIT, MLflow, Comet Opik, Laminar, LangWatch, Latitude, or your own collector. One plugin, parallel fan-out, zero hot-path blocking.

$ hermes plugins install briancaffey/hermes-otel/hermes_otel

Why hermes-otel?​

Hermes Agent is a production agent loop — tools, skills, memory, a gateway, messaging platforms. The moment you ship it, you need to see what it's actually doing: which tools fired, how many tokens the model burned, which turns stalled, which users hit errors.

hermes-otel turns every Hermes lifecycle hook into a properly-nested OpenTelemetry span with the right attribute conventions for the backend you're sending to — no adapter code per vendor. Drop it in, point it at an OTLP endpoint, and the traces show up.

Dual-convention attributes

Emits both gen_ai.* (Langfuse / SigNoz) and llm.token_count.* (Phoenix / OpenInference) so the UI in your chosen backend just works.

Multi-backend fan-out

Send the same span to Phoenix + Langfuse + Jaeger in parallel, each on its own non-blocking worker. One slow collector can't stall the others — or the agent.

Per-turn summary

Root session span gets tool count, tool names, skills used, API-call count, and final status. Dashboards don't need to JOIN across spans to answer "what happened in this turn?"

Non-blocking export

BatchSpanProcessor under the hood: span.end() is a queue push. A slow backend adds zero latency to tool calls or API requests on the hot path.

Privacy mode

Flip capture_previews: false to strip every input/output preview at the source. Metadata (tool names, durations, tokens) still flows.

Orphan-span sweep

Long-abandoned sessions don't leak state: a TTL sweeper finalizes stale root spans with final_status=timed_out so your UI stays clean.

Supported backends​

Phoenix
Arize's OSS LLM observability platform. Local docker or Arize AX cloud. Traces only.
Langfuse
OSS LLM engineering platform. Self-host or cloud. Traces only.
LangSmith
LangChain's tracing platform. Cloud with a free tier. Traces only.
SigNoz
OSS observability platform. Local docker or cloud. Traces + metrics + logs.
Jaeger
The classic distributed-trace UI. Single-container local. Traces only.
Grafana Tempo
Tempo + Grafana stack, OSS or Grafana Cloud. Traces only.
Grafana LGTM
Tempo + Mimir + Loki + Grafana in one image. Local docker. Traces + metrics + logs.
Uptrace
OSS APM on ClickHouse, DSN-style setup. Local docker or cloud. Traces + metrics + logs.
OpenObserve
Single-container observability store with SQL. Local docker or cloud. Traces + metrics + logs.
Parseable
SQL over Parquet with a Traces view. Self-hosted or cloud. Traces + metrics + logs.
Honeycomb
Cloud (US / EU) with a generous free tier. Traces + metrics + logs.
W&B Weave
Weights & Biases' LLM tracing. Cloud, Dedicated Cloud or self-managed. Traces only.
Elastic
Elastic Cloud managed OTLP or a self-hosted EDOT Collector into Elasticsearch + Kibana. Traces + metrics + logs.
OpenLIT
OTel-native LLM observability, two containers, Apache-2.0. Traces + metrics + logs.
MLflow
MLflow 3.6+ tracking server with its GenAI Traces UI, token and cost roll-ups. Self-hosted. Traces only.
Comet Opik
LLM evaluation and tracing platform; threads, span types, usage and cost per span. Self-hosted or Comet cloud. Traces only.
Laminar
Agent-focused tracing with span types, tokens and cost. Self-hosted or Laminar cloud. Traces + logs.
LangWatch
LLM observability and evaluation platform. Self-hosted or cloud. Traces + metrics + logs.
Latitude
Prompt engineering and agent observability on the GenAI semconv. Self-hosted or Latitude cloud. Traces only.
Generic OTLP
Any OTLP/HTTP collector. Drop in an endpoint and go.
Multi-backend
Fan the same spans out to several backends in parallel from one config.yaml.

How this relates to Hermes' built-in telemetry​

Hermes core ships content-free gateway monitoring over OTLP (gateway and cron health, no prompts, tool calls or per-run traces) and a bundled Langfuse-only observability plugin. hermes-otel is the run-level plane, and coexists with both:

Hermes gateway monitoring (core)Bundled Langfuse pluginhermes-otel
ScopeGateway/cron health, content-freePer-run tracesPer-run traces + metrics + logs
BackendsAny OTLP receiverLangfuse only19 backend types plus generic OTLP, fan-out
Coexists with hermes-otelYesYes

The span hierarchy​

agent / cron [root, AGENT]
└── llm.{model} [LLM — input, output, total tokens]
├── api.{model} [LLM — prompt/completion tokens, duration]
│ └── tool.{name} [TOOL — args, result, outcome]
└── api.{model} [LLM — second round-trip, final response]

Each span carries the attributes both Langfuse (gen_ai.usage.input_tokens, gen_ai.content.prompt) and Phoenix (llm.token_count.prompt, input.value) expect — see Attribute conventions.

Where to go next​

🚀 QuickstartInstall + Phoenix in a local Docker container, first trace in under 5 minutes
📦 InstallationInstall into Hermes Agent's venv, optional langsmith extra
🧩 ConceptsHooks, spans, fan-out, how the plugin wires into Hermes
🎯 Pick a backendComparison table, quick picks, decision flowchart
⚙️ Configurationconfig.yaml, env vars, sampling, privacy, batch tuning
🏗️ ArchitectureSpan hierarchy, attribute conventions, turn summaries, orphan sweep
🛠️ ContributingRun the test suite, add a backend, open a PR
📑 ReferenceEvery env var, every config key, every span attribute