Migrating a 17-year-old WebLogic system to microservices with zero downtime
How we moved Customer Vault and Bank Core off a legacy WebLogic platform — by shadowing live traffic with Envoy, proving parity with OpenDiffy, and shifting users over a percentage at a time.
TL;DR
- What: Customer Vault and Bank Core, running on a WebLogic platform for 17 years, rebuilt as microservices.
- Constraint: zero downtime — these services sit on the critical path for payments.
- How: Envoy in front of both stacks, shadowing live requests to the new services; OpenDiffy comparing old and new responses for parity; then percentage-based traffic shifting with instant rollback.
The problem with a 17-year-old system
As Lead Platform Engineer, one of the most consequential projects I led was retiring the WebLogic platform behind two business-critical services: Customer Vault and Bank Core. The platform had been in production for 17 years. It worked — that was the hard part. Nearly two decades of behaviour, edge cases and quiet assumptions were encoded in it, and much of that behaviour was documented nowhere except in the code and in the responses it returned.
Keeping it was getting more expensive every year: an ageing application server, slow and risky releases, a shrinking pool of people who understood it, and an architecture that couldn't scale parts of the system independently. Rewriting it as microservices was the right destination. The real question was how to get there safely.
Why not a big-bang cutover?
The classic approach — build the new system, schedule a maintenance window, switch over and hope — fails on both of our constraints:
- Downtime wasn't acceptable. These services are on the payment path; a maintenance window is a business outage.
- "Hope" isn't a test strategy. No test suite written from today's understanding could capture 17 years of production behaviour. The only complete specification of the legacy system was the legacy system itself.
That second point shaped the whole approach. Instead of asking "does the new system pass our tests?", we asked "does the new system behave exactly like the old one, on real production traffic?" — and we answered it before a single customer request depended on the new code.
The architecture
We put an Envoy proxy in front of both stacks. Every request went through Envoy, which gave us one place to control routing without touching clients or the legacy code.
The migration then ran in four phases:
- ShadowMirror live production requests to the new services. Customers only ever see the legacy response.
- Prove parityUse OpenDiffy to compare legacy and new responses, and fix differences until the diff is clean.
- Shift traffic graduallyMove real users over a small percentage at a time, with rollback a single config change away.
- DecommissionAt 100% and stable, retire the WebLogic platform.
Phase 1 — Shadow traffic with Envoy
Envoy's request_mirror_policies send a copy of each request to a second cluster in a fire-and-forget fashion: the client gets the response from the primary route, and the shadow response is discarded. That let the new services see the full shape of production traffic — real payloads, real volumes, real edge cases — while carrying zero customer risk.
# Envoy route: serve from WebLogic, mirror every request for comparison
routes:
- match: { prefix: "/" }
route:
cluster: weblogic_legacy
request_mirror_policies:
- cluster: diffy_proxy
runtime_fraction:
default_value: { numerator: 100, denominator: HUNDRED }
Simplified example. runtime_fraction lets you start by mirroring a small share of traffic and raise it at runtime.
Two details matter when shadowing anything that touches money or customer data:
- Shadow requests must never cause real side effects. A mirrored write that actually commits is a duplicate transaction. The shadow path has to be isolated from anything customer-visible — its own data store, stubbed downstream calls, or a mode that validates without committing.
- Make shadow traffic identifiable. Envoy appends
-shadowto theHostheader of mirrored requests, which makes them easy to tag in logs and metrics and to exclude from business reporting.
Phase 2 — Proving parity with OpenDiffy
Shadowing tells you the new system responds. It doesn't tell you it responds correctly. For that we used OpenDiffy, the open-source successor to Twitter's Diffy.
Diffy sits behind the mirror and sends each request to three targets:
- Primary — the legacy WebLogic system.
- Secondary — a second legacy instance.
- Candidate — the new microservice.
The secondary is the clever part. Comparing primary with secondary shows which fields differ even between two copies of the same code — timestamps, generated IDs, ordering of unordered collections. Diffy treats those fields as noise and filters them out of the primary-vs-candidate comparison. What's left are real behavioural differences.
# Illustrative OpenDiffy setup
docker run -d --name diffy -p 8880:8880 -p 8888:8888 diffy/diffy \
-candidate=vault-service:8080 \
-master.primary=weblogic-a:7001 \
-master.secondary=weblogic-b:7001 \
-service.protocol=http \
-serviceName=customer-vault \
-proxy.port=:8880 -http.port=:8888 -rootUrl=localhost:8888
Diffy's dashboard groups differences by endpoint and field, which turned "the new system is different somehow" into a concrete, prioritised work list. Each difference fell into one of three buckets:
- A bug in the new service — fix it and redeploy the candidate.
- An undocumented legacy behaviour that clients relied on — replicate it, and write it down this time.
- An intentional, agreed change — exclude it from the comparison explicitly, so it can't hide anything else.
The bar for moving on was a clean diff on real production traffic, endpoint by endpoint — not a passing test suite.
Phase 3 — Percentage-based migration
Once an endpoint showed parity, we started sending it real traffic using Envoy's weighted_clusters. The weight is just configuration, so moving forward — or rolling back — needed no deployment of either application.
# Envoy route: 95% legacy, 5% new — keep mirroring for continued comparison
route:
weighted_clusters:
clusters:
- { name: weblogic_legacy, weight: 95 }
- { name: vault_microservice, weight: 5 }
We ramped in steps, holding at each stage long enough to watch error rates, latency and business metrics before going further:
A few things made the ramp safe:
- Rollback was one config change. If anything looked wrong, the weight went back to the legacy cluster in seconds — no redeploy, no data repair.
- Endpoint by endpoint, not all at once. Routes that reached parity early could move ahead of the rest, so one difficult endpoint didn't block everything else.
- Watch the same metrics on both sides. Error rate, latency percentiles and business-level outcomes, compared between the legacy and new clusters at the same moment.
Phase 4 — Turning off WebLogic
With all traffic on the microservices and the system stable, the WebLogic platform could finally be retired — after 17 years, without customers ever noticing the switch. That's the goal of a migration like this: the best outcome is that nobody outside the team can tell it happened.
Lessons I'd take to the next migration
- Production traffic is the only complete spec of a long-lived system. Design the migration around observing it, not around documents describing it.
- Separate "it runs" from "it's correct". Shadowing proves the first; parity checks prove the second. You need both.
- Invest early in noise filtering. A diff tool that reports thousands of timestamp differences gets ignored. Diffy's primary/secondary trick is what made the signal usable.
- Make every step reversible. When rollback is a config change, small steps are cheap — and small steps are what made zero downtime achievable.
- Treat side effects as the main risk. In payments, a duplicated write is worse than a failed read. Decide how the shadow path stays isolated before mirroring the first request.