title: "Spring 应用可观测性:关联遥测、Grafana 与 MCP 诊断"
source_url: "https://www.youtube.com/watch?v=1LVzd9Wggg0"
author: "Spring I/O"
excerpt: "公共早报 Timo Salm 与 Tiffany Jernigan 将 Spring Boot、Micrometer、OpenTelemetry、Grafana LGTM 和 MCP 串成一套实用的可观测性模型,从关联遥测一路推进到代码级诊断。"
Brief Description
Timo Salm from Broadcom and Tiffany Jernigan from Grafana Labs demystify observability for modern Spring applications by exploring the four core signals: metrics, logs, traces, and continuous profiling. They explain how OpenTelemetry integration has evolved in Spring Boot 3 and 4, demonstrate how Micrometer acts as the primary abstraction for custom observations, and show how to run a local Grafana LGTM stack alongside an AI-assisted Model Context Protocol (MCP) server for deep diagnostics and seamless correlation.
Table of Contents
Introduction and What Is Observability
The Four Signals of Observability: Metrics, Logs, Traces, and Profiles
Correlating Observability Signals
The Role of OpenTelemetry
Observability Evolution in Spring Boot
Custom Instrumentation with Micrometer and Observations API
Clean Architecture Patterns: Separating Observability from Business Logic
Setting Up the Open Source Grafana LGTM Stack
Integrating Model Context Protocol (MCP) with Observability
Live Demo: End-to-End Tracing, Profiling, and AI Exploration
Conclusion and Resources
Introduction and What Is Observability
Tiffany Jernigan: Hello, everybody. Today we are going to be talking about observability for Spring applications. I am Tiffany Jernigan, a Senior Developer Advocate at Grafana Labs. A lot of what we will be showing is based on the open-source side of the ecosystem.
Timo Salm: My name is Timo Salm, based out of Stuttgart. I work for Broadcom, the vendor behind Spring, focusing on helping customers succeed with solutions like Spring, Bitnami, and RabbitMQ.
We will start with the fundamentals of observability, examine what the different observability signals provide, explore what you need to do to get the maximum value out of them, and then walk through demo code showing how easy it is to set up a Grafana stack for local development and testing.
In the past, we were mostly aware of monitoring, which primarily tells you that something is wrong, such as an error appearing in your logs. Observability is much broader because it helps you understand the internal behavior of your applications through signals emitted to the outside world. With a visualization platform like Grafana, you can inspect dashboards, understand application health, and identify where and why issues occur.
With agile development and modern practices, the operational responsibility for applications is increasingly in the hands of developers. That does not mean developers should have to SSH into Kubernetes clusters to figure out problems. Observability provides the necessary visibility so developers can inspect and diagnose issues directly from telemetry data.
The Four Signals of Observability: Metrics, Logs, Traces, and Profiles
Timo Salm: There are three plus one core signals in modern observability. The first is metrics, which represent numeric, computed measurements aggregated over time, ranging from JVM memory consumption to critical business KPIs. Business metrics are an area where many applications still fall short.
The second signal is logs, which help you capture discrete events and errors. The third signal is distributed traces, which are essential for modern distributed architectures to trace request lifecycles across services, identify root causes of failures, and locate latency bottlenecks.
The newest addition is continuous profiling, which helps you pinpoint performance bottlenecks down to specific lines of code. While a trace might tell you that a particular method took twenty seconds, a profile reveals exactly which code path or instruction consumed CPU cycles or memory allocations during that execution.
Out of the box, frameworks like Spring Boot provide automatic metrics for JVM resource consumption, and extensions like Spring AI supply token usage statistics, but you can also add custom instrumentation for deeper visibility.
Tiffany Jernigan: To think about metrics conceptually, imagine tracking the temperature in a room over time, or monitoring CPU and memory usage over intervals.
Logs represent discrete events with contextual information. You might log when an application starts up, record business logic events, or audit access when a user signs in and grants permissions. Logs support various severities, including debug, info, warning, and error, and can be structured as plain text or JSON for downstream processing.
Distributed traces follow the full journey of a request across your infrastructure. When a user interacts with a web application, the request may travel through an API gateway, multiple microservices, external dependencies, and databases. The entire end-to-end journey is a trace, while individual units of work within that path, such as an HTTP GET or a database query, are called spans. You can customize spans to capture application-specific execution paths even within a monolith.
Profiling captures runtime resource usage across functions, highlighting unexpected memory allocations, garbage collection pressure, or CPU-intensive methods.
Timo Salm: There is a distinct difference between traditional profiling used during isolated testing and continuous profiling. Traditional profiling tools capture extensive detail but impose significant overhead. Continuous profiling minimizes resource consumption and overhead, making it safe to run continuously in production to monitor performance trends over time.
Tiffany Jernigan: Continuous profiling allows you to catch intermittent anomalies and compare performance between different deployment versions or time windows. Profiling data is commonly represented as top tables or flame graphs, a visualization popularized by Brendan Gregg.
Correlating Observability Signals
Timo Salm: Having individual signals is useful, but the real power of observability comes from correlating them. When your telemetry signals are connected, you can start with an error log, jump straight into the corresponding distributed trace to pinpoint the failing downstream dependency, inspect relevant metrics, and verify the root cause without guessing.
Tiffany Jernigan: A helpful mental model is that metrics indicate whether something is happening, logs tell you what is happening, traces show where it is happening across your systems, and profiles help you understand how to fix code-level inefficiencies.
Correlating these signals through shared identifiers like trace IDs prevents developers from having to manually cross-reference disconnected tools.
The Role of OpenTelemetry
Tiffany Jernigan: In the past, engineering teams relied on siloed tools and proprietary agents for different telemetry types, such as maintaining separate setups for Elasticsearch, Kibana, Prometheus, or Jaeger. Switching a vendor or backend required rewriting instrumentation across codebases.
OpenTelemetry emerged within the Cloud Native Computing Foundation (CNCF) to provide an open, vendor-neutral standard for collecting and exporting telemetry data. Today, OpenTelemetry Protocol (OTLP) specifications are stable for traces, metrics, and logs, with profiling specifications progressing through standardization.
Observability Evolution in Spring Boot
Timo Salm: Spring Boot is known for high-level abstractions and convention-over-configuration defaults designed to make applications production-ready with minimal setup.
In Spring Boot 3, Micrometer serves as the core facade on top of various observability frameworks and protocols. Because Micrometer existed before OpenTelemetry became the dominant industry standard, it provided a unified API that abstracted away lower-level protocol details.
Developers using Spring Boot have historically had three primary options for OpenTelemetry:
The OpenTelemetry Java Agent: A zero-code solution that instruments code at runtime, commonly injected via container build mechanisms such as Cloud Native Buildpacks.
Community Starters: Community-driven OpenTelemetry starters, which lack official enterprise support.
Micrometer with OpenTelemetry Exporters: The recommended approach using Micrometer's native abstractions combined with OTLP exporters.
Spring Boot 4 expands native support for OpenTelemetry dependencies and starters. Spring Boot 4.1 introduces further refinements, including standard observation conventions across components like database drivers and built-in support for metric exemplars. Exemplars link specific metric data points directly to trace IDs, enabling immediate navigation from an anomaly on a metric graph to the corresponding trace.
Custom Instrumentation with Micrometer and Observations API
Timo Salm: While framework auto-configuration captures external network boundaries and database interactions, capturing domain-level business context requires explicit instrumentation. Micrometer's Observation API allows developers to track business events and custom execution paths cleanly.
You can instrument code declaratively using annotations such as @Timed or @Counted, or imperatively by wrapping code blocks using the ObservationRegistry.
Tiffany Jernigan: When adding attributes to an observation, you can specify low-cardinality or high-cardinality tags. High-cardinality attributes, such as specific SKU IDs or order quantities, are attached to span data within distributed traces rather than polluting metric aggregations.
Clean Architecture Patterns: Separating Observability from Business Logic
Timo Salm: Injecting observation logic directly into core business methods can introduce clutter and violate separation of concerns. To keep business logic clean and readable, you can decouple telemetry using established design patterns:
Decorator Pattern: Wrap domain services with a decorator implementation that implements the same interface, handles observation logic, and delegates execution to the core bean.
Aspect-Oriented Programming (AOP): Use Spring Aspects to intercept method executions and update counters, timers, or distribution summaries via the
MeterRegistryafter methods complete.
This ensures the domain layer remains focused purely on business rules while telemetry is managed independently.
Setting Up the Open Source Grafana LGTM Stack
Tiffany Jernigan: To visualize and analyze these telemetry signals, you can run the open-source Grafana LGTM stack, which includes:
Loki: Log aggregation backend.
Grafana: Visualization and unified dashboarding interface.
Tempo: Distributed tracing backend.
Mimir / Prometheus: Metric storage and querying.
Pyroscope: Continuous profiling storage.
The OpenTelemetry Collector, or Grafana Alloy, acts as the ingestion pipeline receiving OTLP data over gRPC (port 4317) or HTTP (port 4318) and routing it to the respective backend storage engines.
Timo Salm: Grafana provides pre-packaged all-in-one container images for local testing. Combined with Docker Compose and Cloud Native Buildpacks (./mvnw spring-boot:build-image), developers can spin up a fully instrumented environment locally with pre-provisioned dashboards without writing custom Dockerfiles.
Integrating Model Context Protocol (MCP) with Observability
Tiffany Jernigan: The widespread adoption of AI coding assistants has increased the volume of code deployed to production, making testing and observability more critical than ever. The Model Context Protocol (MCP), open-sourced by Anthropic, acts as a standardized interface allowing AI models to interact securely with external development tools and services.
Grafana provides an open-source MCP server that exposes tools for querying Loki, Tempo, Prometheus, and Grafana dashboards directly to AI assistants like Claude or Cursor.
Timo Salm: By exposing local or remote telemetry endpoints to an MCP client, developers can ask an AI assistant in natural language to inspect failing services, summarize recent error traces, or generate tailored Grafana dashboard JSON configurations automatically based on live telemetry data.
Live Demo: End-to-End Tracing, Profiling, and AI Exploration
Tiffany Jernigan: In the demo application, an order service processes transactions and calls a mystery box service powered by Spring AI and OpenAI's GPT-4o-mini.
Within Grafana, looking at a distributed trace reveals the full call hierarchy:
The inbound HTTP request hits the order service.
The order service delegates item processing and issues downstream requests.
The mystery box service invokes Spring AI, exposing span attributes for LLM token usage and execution latency.
Custom span attributes show high-cardinality metadata such as product SKU names and ordered quantities.
Using configured data source links, you can jump directly from a span into Loki to view contextual application logs or navigate to Pyroscope flame graphs to inspect JVM method execution profiles.
Timo Salm: When an MCP server is configured within an IDE or chat interface, you can query the system conversationally to list active services, inspect system health, and diagnose performance regressions using the connected Grafana data sources.
Injecting trace IDs into downstream HTTP response headers also ensures that frontend clients or downstream consumers can immediately report the trace ID for rapid troubleshooting.
Conclusion and Resources
Tiffany Jernigan: Observability allows teams to understand distributed system behavior through correlated metrics, logs, traces, and profiles. Using Spring Boot's native Micrometer abstractions combined with OpenTelemetry standards provides a flexible, production-ready foundation.
All sample code, Docker configurations, and demonstration repositories are available on GitHub for developers looking to implement full-stack observability with Spring Boot and Grafana.