Centralize structured events at three boundaries: agent run, model attempt, and tool call. The deciding constraint is signal quality, not ingestion speed. A small team can ship searchable application logs quickly, but those logs become expensive noise unless every record answers one operational question: where did this run wait, fail, or spend its budget?
Short answer: emit newline-delimited JSON to standard output, let the runtime forward it to one searchable store, and give each boundary a stable correlation ID. Record durations and usage at completion, retain a small failure-safe start event, and keep prompts, responses, secrets, and user contact data out by default. This pattern works across Python, JavaScript, and Ruby services without putting a logging client on the request path.
How should a small SaaS add centralized application logs?
Treat the event schema as an architecture decision record, not a bag of fields. The system has three invariants. First, one run_id follows an agent loop across queues and services. Second, each model attempt and tool call has its own child ID, because retries must not overwrite the work they caused. Third, completion events carry measured duration, outcome, and usage from the component that observed them.
Those rules create a useful failure boundary. A process can die after emitting model_attempt_started but before model_attempt_finished; the unmatched start then means "unknown completion," not zero latency or zero cost. Be precise. An alerting pipeline that quietly converts missing values to zero will make a delivery gap look healthy, much like an OTP provider timeout that gets counted as a successful send because the enqueue call returned.
The minimum event set is small: agent_run_finished, model_attempt_started, model_attempt_finished, tool_call_finished, and agent_run_failed. The start record exists only where abrupt termination would otherwise erase evidence. Routine successful tool calls generally need one completion record.
Five event types are enough to begin.
Use allowlists for fields. Hashing an email address does not automatically make it safe: stable hashes remain correlatable, and low-entropy identifiers may be guessed. For an agent that drafts developer documentation, useful dimensions include operation name, model family, tool name, attempt number, outcome, duration bucket, and token counts reported by the model interface. Raw prompts and tool payloads belong in a separately governed diagnostic workflow, if they are retained at all.
The decision table
There are several credible transport shapes. The right choice depends on how much loss and request-path coupling the service can tolerate, not on the language used to build it.
| Option | Request-path risk | Failure behavior | Operational load | Valid fit |
|---|---|---|---|---|
| Structured stdout plus runtime collector | Low; local write only | Collector or platform applies buffering and retry policy | Low | Small SaaS services already running under a managed runtime or container scheduler |
| In-process network handler | Higher; DNS, TLS, backpressure, and retries enter the application | Must bound queues and define drop behavior | Medium | Environments without a collector where limited loss is explicitly acceptable |
| Local durable agent | Low after local handoff | Can survive brief network outages, subject to disk limits | Higher | Regulated or high-volume systems with a clear host-management owner |
| Database table | Adds database work to application traffic | Competes with product queries and requires retention jobs | Deceptively high | Short-lived prototypes with tiny volume and an explicit migration date |
For a small backend with no dedicated operations team, structured stdout is usually the narrowest failure surface. This is an architectural preference, not a product recommendation. The runtime owns collection; the application owns meaning. Keep those responsibilities separate so a logging outage cannot hold an agent response, password-reset email, or OTP request open.
It also crosses stack boundaries cleanly. A FastAPI service can use Python logging, a Node.js worker can emit the same JSON keys, and a Rails application can map its logger output to that shared envelope. The centralized collector does not need to understand three application frameworks; it needs valid records, timestamps, and the agreed identifiers. That matters when a small SaaS has no DevOps specialist available to maintain a different shipping path for every runtime.
The store still needs a service-level objective. Decide how late an event may arrive before an operator considers search incomplete, then test that delay. Also set retention by event class rather than treating every debug line as equally valuable. Failure summaries may deserve longer retention than verbose success details, provided the policy matches privacy and compliance obligations.
Put measurement on the critical path, not transport
The application should measure work synchronously and export asynchronously. Python's standard logging API can produce a stable schema without binding business logic to a remote endpoint. This example records one model attempt; the injected logger can write JSON to stdout, while the caller supplies identifiers created at the relevant boundaries.
import logging
import time
from collections.abc import Callable
from typing import Any
logger = logging.getLogger("agent.telemetry")
def invoke_model(
call: Callable[[], dict[str, Any]],
*,
run_id: str,
attempt_id: str,
model_family: str,
) -> dict[str, Any]:
started_ns = time.monotonic_ns()
logger.info(
"model_attempt_started",
extra={
"event": "model_attempt_started",
"run_id": run_id,
"attempt_id": attempt_id,
"model_family": model_family,
},
)
try:
result = call()
except Exception as exc:
logger.exception(
"model_attempt_finished",
extra={
"event": "model_attempt_finished",
"run_id": run_id,
"attempt_id": attempt_id,
"model_family": model_family,
"outcome": "error",
"error_type": type(exc).__name__,
"duration_ms": (time.monotonic_ns() - started_ns) // 1_000_000,
},
)
raise
usage = result.get("usage", {})
logger.info(
"model_attempt_finished",
extra={
"event": "model_attempt_finished",
"run_id": run_id,
"attempt_id": attempt_id,
"model_family": model_family,
"outcome": "ok",
"duration_ms": (time.monotonic_ns() - started_ns) // 1_000_000,
"input_tokens": usage.get("input_tokens"),
"output_tokens": usage.get("output_tokens"),
},
)
return result
time.monotonic_ns() is appropriate for elapsed time because it cannot go backward; wall-clock timestamps should still be added by the log formatter for cross-service search. Do not compute spend from a stale price constant inside this function. Preserve the provider-reported usage, attach a versioned rate-card reference during a downstream enrichment step, and allow the cost to remain unknown when usage is absent. Accuracy beats a comforting zero.
The exception record intentionally contains an error class but not str(exc). Exception messages from HTTP clients and tool adapters can include URLs, query strings, response fragments, or recipient data. A reviewed sanitizer may later admit selected messages, but the default schema should be safe enough for broad operational access.
Unknown stays unknown.
Search from a question, then control cardinality
Start with three saved investigations rather than a dashboard wall: slow agent runs by phase, failed runs grouped by bounded error type, and token usage by operation and model family. These map to latency, errors, and resource consumption, three of the four golden signals described in the Google SRE guidance. Traffic can be represented by run counts. Saturation usually needs runtime metrics such as queue depth or worker utilization; forcing it into application logs produces awkward, duplicated samples.
A useful latency investigation joins completion events by run_id, compares summed model and tool durations with total run duration, and leaves the remainder visible as orchestration overhead. Retry attempts stay separate. For cost attribution, aggregate recorded input and output tokens first, then apply the rate-card version used for the reporting period. This makes recalculation possible and stops a pricing update from changing historical meaning silently.
Cardinality is the trap. run_id and attempt_id are search keys, but they should not become metric labels or default group-by dimensions. Tool names and model families should come from bounded registries. URLs should be normalized to route templates. Error messages, prompt excerpts, stack traces, and recipient identifiers are payloads, not dimensions.
Set a daily event budget from expected traffic and fan-out. For example, a run with two model attempts and three tool calls yields at least eight records under this schema: one run completion, two attempt starts, two attempt completions, and three tool completions. That number is not a benchmark or a promise; it is arithmetic the team can substitute into its own peak-run forecast. Sample repetitive successes only after preserving aggregate counts and all failures.
Then exercise the ugly paths: a killed worker after a start record, a tool timeout followed by retry, a queue redelivery with the same run ID, a malformed usage object, and a collector outage. Follow one concrete run through all five cases. The first attempt starts and the worker dies, so search shows an unmatched attempt rather than a fabricated duration. The queue redelivers the run, and the second worker creates a new attempt ID under the original run ID. A tool times out; its completion record says error, while the retry gets another child ID. The model response then omits usage, which must remain null. Finally, stop the collector and confirm that the application keeps serving while the runtime's documented buffering or drop policy takes effect. The acceptance test is not merely "logs appeared." Search must distinguish incomplete from completed attempts, retries from duplicates, and unknown usage from zero usage.
Test the gaps.
Rejected option and the case where it works
This decision rejects direct synchronous delivery from each request handler to a remote log API. It expands the blast radius: connection setup, rate limits, backpressure, and retry queues now compete with user work. That is a poor default for a small team trying to diagnose agent latency without creating another source of it.
The stdout approach has a real limitation: delivery guarantees belong to the runtime. If the platform can discard output immediately when an instance terminates, application code cannot repair that gap after the fact. The trade-off is acceptable only when the documented buffering behavior and measured loss window match the service's evidence requirements.
The option remains valid in constrained serverless environments where stdout collection cannot meet delivery requirements and no sidecar or host agent is available. In that case, use a bounded in-memory queue, batch sends, short timeouts, and an explicit overflow policy. Measure dropped records locally. Never retry without a cap, and never make successful agent work wait indefinitely for telemetry.
No transport is lossless by declaration.
A custom appender or handler can enforce those rules, but it becomes production infrastructure. The Logback appender documentation illustrates the lifecycle and synchronization concerns that such extensions must handle. Other language runtimes have different APIs, yet the operational questions stay the same: who buffers, what blocks, what is dropped, and how shutdown flushes are bounded.
The final decision is deliberately plain. Centralize transport outside the application, standardize meaning inside it, and spend the event budget on reconstructing boundaries. That gives a small SaaS team searchable evidence for agent latency and usage while keeping noise, sensitive data, and telemetry failure away from the product path.
References
- https://sre.google/sre-book/monitoring-distributed-systems/
- https://docs.python.org/3/library/logging.html
- https://docs.python.org/3/library/time.html#time.monotonic_ns
- https://opentelemetry.io/docs/specs/otel/logs/data-model/
- https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html
- https://logback.qos.ch/manual/appenders.html
Top comments (0)