Most teams do not fail at microservices because they picked the wrong framework. They fail because they split a system before they understood it, then discovered that a network call is a very different thing from a method call. This is the set of practices we reach for when a Java system built on Spring Boot has to keep running under real load.
1. Start with boundaries, not frameworks
A service should own a business capability and the data behind it. If two services must always be deployed together, or share a database table, they are one service wearing a disguise.
- Draw boundaries around capabilities such as ordering, billing or inventory, not around technical layers.
- One service, one datastore. Other services get at the data through an API or events, never through the database.
- Earn the split. A well-structured modular monolith is a perfectly good starting point. Split a module out when it needs to scale, deploy or be owned independently.
Rule of thumbIf you cannot describe what a service owns in one sentence, the boundary is probably wrong.
2. A sensible service skeleton
Keep every service boring and consistent so that engineers can move between them. On a current stack that means Java 21, Spring Boot 3.x, and a small set of conventions applied everywhere: externalised configuration, health endpoints, structured logging and a standard error format.
Java 21 also gives you virtual threads. For services that spend most of their time waiting on I/O, switching them on is a single property and often removes the need for reactive code:
# application.yml
spring:
threads:
virtual:
enabled: true
server:
shutdown: graceful
management:
endpoint:
health:
probes:
enabled: true
3. Treat every remote call as something that will fail
Synchronous calls between services are the most common source of cascading failure. Three defences cover most of the risk: timeouts so a slow dependency cannot hold your threads hostage, retries only for idempotent operations, and a circuit breaker so a failing dependency is given room to recover.
@Service
class InventoryClient {
private final RestClient http;
InventoryClient(RestClient.Builder builder,
@Value("${inventory.url}") String baseUrl) {
var factory = new JdkClientHttpRequestFactory(
HttpClient.newBuilder()
.connectTimeout(Duration.ofMillis(500))
.build());
factory.setReadTimeout(Duration.ofSeconds(2));
this.http = builder.baseUrl(baseUrl).requestFactory(factory).build();
}
@CircuitBreaker(name = "inventory", fallbackMethod = "unknownStock")
@Retry(name = "inventory") // safe: this is a GET
StockLevel stockFor(String sku) {
return http.get().uri("/stock/{sku}", sku)
.retrieve().body(StockLevel.class);
}
StockLevel unknownStock(String sku, Throwable cause) {
return StockLevel.unknown(sku); // degrade, do not crash
}
}
Resilience4j annotations with Spring Boot 3 (resilience4j-spring-boot3)
Decide up front what a sensible fallback is. For a stock lookup, showing "availability unknown" is usually better than failing the whole page.
4. Use events, and publish them safely with an outbox
Where a business action should trigger work in other services, publish an event rather than calling them one by one. The classic trap is the dual write: you save to the database, then publish to the broker, and the process dies in between. The database says the order exists, and nobody else ever hears about it.
The transactional outbox removes the gap. You write the event to an outbox table in the same transaction as the business change, and a separate relay publishes it afterwards.
@Transactional
public Order place(NewOrder cmd) {
Order order = orders.save(Order.from(cmd));
outbox.save(OutboxEvent.of("order.placed", order.getId(), toJson(order)));
return order; // one transaction, two rows
}
@Scheduled(fixedDelay = 1000)
void relay() {
for (OutboxEvent e : outbox.findUnsentBatch()) {
broker.publish(e.getTopic(), e.getKey(), e.getPayload());
e.markSent();
}
}
This gives you at-least-once delivery, which means consumers must be idempotent. Give every event a unique id and have consumers ignore ids they have already processed.
5. Accept that consistency is now your job
Once data lives in several services you cannot wrap a business operation in one database transaction. A saga breaks the operation into local steps, each with a compensating action: reserve stock, take payment, and if payment fails, release the stock. Keep sagas short, make every step idempotent, and design the compensations at the same time as the happy path.
6. Build observability in from day one
With one application you read one log file. With twenty you need to follow a request across all of them.
- Correlation. Propagate a trace id on every call and every event, and put it in every log line.
- Metrics. Micrometer ships with Spring Boot. Track the four golden signals: latency, traffic, errors and saturation.
- Tracing. OpenTelemetry gives you a picture of where a slow request actually spent its time.
- Structured logs. Log JSON, not sentences, so the platform can search it.
7. Deploy to Kubernetes like you mean it
A container that starts is not the same as a service that is ready. Give Kubernetes the signals it needs to run your service safely:
- Liveness and readiness probes backed by the Actuator health groups, so traffic only reaches instances that are actually ready.
- Graceful shutdown so in-flight requests finish before a pod is replaced.
- Resource requests and limits set from measurement, with the JVM memory settings aligned to the container limit.
- Configuration and secrets injected from the platform, never baked into the image.
8. Test the seams
Unit tests will not tell you that two services disagree about a field name. Add contract tests (Pact or Spring Cloud Contract) so a provider cannot break a consumer without a failing build, and use Testcontainers to run integration tests against a real database and broker instead of mocks.
A short checklist
- Each service owns a capability and its data.
- Every remote call has a timeout, and a circuit breaker where it matters.
- Events are published through an outbox and consumed idempotently.
- Traces, metrics and structured logs exist before the first release.
- Probes, graceful shutdown and resource limits are configured.
- Contracts are tested in the pipeline.
None of this is glamorous, and that is the point. Reliable microservices are mostly a matter of applying a small number of unglamorous habits consistently.