Best Practices for IT Modernization with Observability
The best practices for IT modernization with observability start before migration: define the user outcomes that matter, measure the existing service, instrument the old and new paths, and use that evidence to control each rollout. Logs, metrics and traces become useful when they answer whether a change improved reliability, performance and operating cost.
Modernization might mean updating a runtime, replacing a database, moving selected workloads to the cloud or separating part of a legacy application. It does not require converting everything into microservices. I would choose the smallest meaningful change whose result the team can measure and operate confidently.
Start with a business journey and a measurable decision
Choose a journey such as submitting an order, retrieving an account balance or completing a nightly report. Identify who depends on it, what successful completion means and how long users can reasonably wait.
Then turn the modernization goal into a decision. “Move this application to the cloud” describes an implementation. “Keep order submission reliable during peak demand while reducing operating effort” describes an outcome that telemetry can help evaluate.
Useful starting question
Can customers complete the same task successfully and within the agreed response time on the new path?
Evidence to collect
Eligible request counts, successful outcomes, latency, failed dependencies and the cost of delivering that workload.
Observability helps teams investigate system behaviour using emitted data. Monitoring known thresholds remains valuable, but modernization also introduces unfamiliar failure modes. A useful setup must support investigation beyond a predetermined dashboard.
For background on the relationship between signals and operational context, this observability principles overview provides a broader introduction. The implementation should still begin with your service’s decisions and owners.
Build a baseline before changing the architecture
Record the existing service under representative conditions: ordinary demand, important peaks, scheduled batch work and relevant failure scenarios. A quiet afternoon is a weak baseline for a month-end reporting system.
Capture throughput, success rate, latency distribution, saturation and dependency behaviour. Keep workload details alongside the measurements so that a lighter query mix or smaller dataset is not mistaken for a successful optimization.
Use comparable definitions across environments. If the old system measures the complete user request but the new system measures only an internal API call, their latency numbers do not describe the same experience. Likewise, a successful HTTP response might still contain a business-level failure.
I would retain deployment markers, configuration versions and the measurement window with each comparison. For latency, examine percentiles such as p95 alongside counts and distributions; an average can hide a slow minority. Avoid averaging separately calculated percentiles as though that produces a valid combined percentile.
Connect telemetry across legacy and modern components
A migrated front end may still depend on a legacy adapter, shared database, identity service or message queue. Instrument the boundaries between them before declaring the new component observable.
Use metrics to detect a change in volume, errors or latency; traces to investigate the path of a request; and logs for relevant event detail. Correlate those signals through consistent service identity and request context where the instrumentation supports it.
OpenTelemetry supplies vendor-neutral instrumentation and telemetry collection mechanisms. It is not, by itself, a complete storage and visualization backend. You still need somewhere to retain, query and use the exported signals.
Propagate trace context across supported service calls and messaging boundaries. Where an older component cannot participate, record the gap and use boundary timings or a carefully designed correlation identifier. Do not draw an apparently complete dependency map from incomplete instrumentation.
Automatic instrumentation can expose common framework operations, while business outcomes may need explicit instrumentation. For example, a span that shows a database call finished does not prove the order was accepted correctly.
Standardize names without creating a telemetry explosion
Agree service names, environment labels, version identifiers, metric units and ownership before expanding instrumentation. OpenTelemetry semantic conventions provide shared definitions that help data from different libraries and platforms remain understandable.
Document the convention versions and compatibility decisions you adopt. A library update that changes an attribute name can break a dashboard or release query even when the application itself still works. Treat telemetry changes as part of the release review.
Keep metric dimensions bounded. Prometheus identifies each unique combination of label values as a distinct time series. Adding individual user IDs, email addresses or arbitrary request paths as labels can create an unbounded number of series.
I would use a route template such as an order endpoint pattern where supported, rather than every individual order URL. Keep request-specific investigation in appropriately controlled logs or traces, and avoid collecting identifiers without a clear operational need.
Set service objectives that guide rollout decisions
A service-level indicator, or SLI, measures behaviour such as the proportion of eligible requests completed successfully. A service-level objective, or SLO, sets the target over a defined window. Decide which requests count, how success is classified and what happens when measurements are missing.
In the hypothetical example, 25 failed requests out of 100,000 leave the success rate at 99.975%. That is above the 99.9% objective for the stated completed window. The calculation illustrates the definition; it is not a reason to accept every failure or choose that target for every application.
For rollout decisions, also inspect the affected cohort. A small new deployment can fail badly while the aggregate service remains within its objective because most traffic still uses the old version.
Keep correctness separate from transport success. Duplicate orders, missing records or a report containing wrong totals can be serious regressions even when latency and HTTP status codes look healthy. Define reconciliation checks where the change affects stored data or business results.
Release in stages with explicit pass, hold and recovery rules
A canary rollout exposes a limited portion of traffic or infrastructure to a change while the remaining portion provides a control. Use comparable workloads and account for differences such as region, customer type and cache state.
Agree acceptance thresholds, minimum useful traffic and an observation period before starting. There is no universal percentage or five-minute duration that proves a migration safe; infrequent batch jobs and low-volume paths may need additional evidence.
Annotate releases in the telemetry and retain a way to identify the deployed version. Compare user outcomes, resource behaviour and data correctness. A new version that is faster only because it rejects more work has not met the intended goal.
Prepare recovery before expansion. Application rollback may not reverse a database schema change, an external side effect or a data transformation. Test compatibility and define whether recovery requires restoring traffic, disabling a feature or making a corrective change.
Connectivity changes need the same discipline. This network policy workflow explains why checking the installed configuration and actual traffic path matters after deployment.
Alert on actionable service impact
Give each urgent alert an owner, a reason to act and a short runbook. A CPU threshold can help diagnosis, but a page should make clear what service is threatened and what the responder can do.
Error-budget burn rate expresses how quickly failures consume the allowance implied by an SLO. Google’s SRE guidance describes using multiple windows to detect both rapid and sustained consumption. Tune that approach to the service rather than copying example thresholds without considering request volume.
Low-traffic services need particular care: one failed request can produce a dramatic percentage. Combine an appropriate observation window with request counts and other evidence. Synthetic checks can test selected paths, but their coverage is not equivalent to the full production workload.
I would review alert quality after a pilot incident: did the right person receive it, could they identify the affected journey, and did the runbook help? A larger number of alerts is not evidence of better coverage.
Control telemetry cost and sensitive data
Choose retention and collection policies by investigative value. Metrics, detailed logs and traces need not have identical retention periods. Estimate normal and peak ingestion, then measure the actual cost and application overhead during the pilot.
Head sampling selects traces early; it cannot guarantee keeping every trace that later encounters an error. Tail sampling can consider later information, including latency or failure, but requires processing capacity and enough trace data to make its decision. It cannot recover spans already discarded upstream.
Do not calculate the overall failure rate from a deliberately error-biased trace sample without appropriate statistical treatment. Use suitable request counters for the SLI and sampled traces to investigate individual paths.
Exclude secrets, access tokens and unnecessary personal data from telemetry. Review instrumentation output as well as application logging, and apply appropriate access, filtering and retention controls. OpenTelemetry does not automatically know which fields are sensitive in your business.
Measure the cost of the whole service, including retained legacy systems, parallel operation and observability. The mainframe cost guide explains why a lower cloud bill alone may not establish a cheaper modernization outcome.
Monitor the telemetry pipeline itself
An empty dashboard can mean no errors, no traffic or no telemetry. Those situations require different responses. Track whether expected data is arriving and whether collectors and exporters are keeping up.
OpenTelemetry Collector internal telemetry exposes information about queue occupancy, enqueue failures, refused data and export behaviour. Use the metrics supported by your deployed version to detect pressure or loss rather than assuming the monitoring system is always healthy.
During a controlled test, interrupt a telemetry destination or simulate a constrained collector. Verify that operators can detect the gap and that buffering, retries and application behaviour match the intended design. Keep the test scoped to an environment where its effects are understood.
My release rule would be explicit: insufficient or stale telemetry prevents automatic expansion. It should trigger investigation, not silently convert an unknown result into a successful rollout.
Use a small pilot to establish the operating model
Start with one important journey and the team responsible for it. Build the baseline, instrument the required boundaries, establish objectives and rehearse one failure and recovery scenario. Then review the evidence before extending the approach.
| Stage | Required evidence | Decision |
|---|---|---|
| Before change | Baseline, dependency gaps, owners and agreed objectives. | Is the comparison meaningful? |
| Limited rollout | Valid telemetry, cohort results and correctness checks. | Expand, hold or recover? |
| After rollout | Representative workload results and operating cost. | Did the change meet its goal? |
| Retirement | Confirmed consumers, recovery needs and retained obligations. | Can the old component be removed? |
Keep responsibility with the people who operate the service. A platform team can provide collection standards and shared tooling, while application owners define meaningful outcomes and response procedures. Neither group should assume the other is covering an unassigned dependency.
Frequently asked questions
Do we need microservices to use observability?
No. A monolith, virtual machine or legacy application can benefit from useful metrics, logs and traces. Choose architecture changes for the workload’s needs, not to qualify for a monitoring approach.
Should we replace all monitoring tools first?
Not necessarily. Begin with the visibility gaps affecting the pilot. Consolidate tools when it solves a demonstrated integration or operating problem, and check that required historical data remains available.
Can AI determine the root cause automatically?
It can assist an investigation, but treat its explanation as a hypothesis to verify. Correlation with a deployment or a noisy component does not by itself prove causation or justify an automatic production change.
Make the next migration an evidence-based decision
Choose one service change and write down the outcome, the measurements, the rollout criteria and the recovery owner. If the team cannot distinguish success, failure and missing evidence, improve that visibility before expanding the migration. That gives observability a concrete job in modernization: helping people decide what to change next and when to stop.