Portfolio Manifest — Last updated 2026-09-09
Backend & Distributed Systems Engineer
An original implementation of the Distributed Saga pattern across six independent Spring Boot / gRPC microservices — a JWT-guarded API gateway in front of order placement, payment settlement, and restaurant fulfillment, coordinated through Kafka with compensating transactions instead of a shared database transaction. Built and verified one service at a time against real Postgres and Kafka infrastructure — not mocked past the point where mocking stops proving anything.
A whiteboard topic, built for real
Saga orchestration is a common interview-whiteboard topic and an uncommon thing to actually build end to end: compensating transactions, event ordering, idempotency, and the failure modes that only show up once services genuinely run independently. This isn't a fork of any existing project — the module boundaries and general shape of the problem (order → payment → fulfillment, with compensation on failure) are common territory for this class of system, but the code, design decisions, and tradeoffs recorded here are this repo's own, built one service at a time with the reasoning written down as it happened.
Outbox pattern, compensating transactions, wired end-to-end behind a real security perimeter
api-gateway-service fronts the chain — JWT-guarded
routing and a Resilience4j circuit breaker sit in front of everything
below. order-service creates an order and stages an
outbox record in the same transaction; a poller ships it to Kafka.
payment-service consumes it, settles payment, and
publishes its own event. restaurant-service consumes
that, allocates inventory, and makes the saga's real decision —
approve or reject — publishing back to both
order-service and payment-service so a
rejection compensates the payment and cancels the order. No shared
database transaction anywhere in the chain.
spring-kafka directly instead of a Spring Cloud
Stream binder — this project only ever targets Kafka, so the
broker-swappability abstraction wasn't worth the config overhead.
UUID ids and BigDecimal money instead of
raw strings and doubles — floating point can't exactly represent
decimal fractions, a real correctness concern once money's
involved. Reasoning for both recorded in
docs/architecture.md.
Every state-changing method follows the same cheap-check-before-load
pattern: an existsBy… query first, then a status
re-check after load, so a duplicate Kafka delivery — the normal
case for at-least-once delivery, not an edge case — can't double-
apply an order confirmation, a payment, or a compensation.
Found via actually running the thing, not reading the source twice
Every entry below was caught by a real, failing run — a test, a CI job, or a live boot against real infrastructure — not by static reading, and every fix was re-verified with another real run before being called done.
spring-grpc's BOM was silently pulling
protobuf-java below what user-contract's
generated code needs, on user-service's actual
runtime classpath — not just tests — since the module's first
phase. Undetected until a test finally constructed a generated
message at runtime. Fixed with an explicit version pin.
contextLoads) failed
outright — Spring Framework's own bundled ASM can't parse JDK 25
class files during component scanning at all, deeper than an
earlier Gradle-plugin-only version of the same issue. Checking
user-service showed it had the identical latent
risk, never caught because no test there had ever booted a real
ApplicationContext. Fixed by bumping both modules to
Boot 3.5.16.
Permission
denied, exit 126) — gradlew had been
committed from Windows without its executable bit, so git stored
mode 100644 instead of 100755. Fixed
with git update-index --chmod=+x, confirmed on a
second real run.
JwtPerimeterGuardGatewayFilterFactoryTest
and actually running it — never assumed — showed two of its
three cases asserting an HTTP response status the filter itself
never sets; that translation only happens through
GlobalExceptionHandler, a layer absent from an
isolated filter unit test. Failed immediately with
onError instead of the expected
onComplete. Rewritten to assert the filter's real,
provable contract instead.
OrderController logging a
whole request record whose itemCode field is
free-form client input with no character restriction —
\r/\n bytes could forge fake log
lines. Fixed by logging fields individually and stripping
control characters from the one field that actually needed it.
PasswordEncoder with a literal
placeholder hash — so the real bcrypt encode()
path had never run once, in production code or tests. Added
a real encode/matches round-trip test.
Login returned three distinguishable gRPC
statuses depending on unknown-email vs. wrong-password vs.
inactive-account, letting a caller enumerate registered
emails — and their active status — without ever guessing a
password. Collapsed to one generic UNAUTHENTICATED
response after weighing the UX tradeoff.
api-gateway-service already
caches its JWTVerifier once at construction for
performance; this file rebuilt one on every single
verifyToken() call. Matched the established
pattern — zero behavior change, existing tests passed
unchanged as proof.
user-service
had no @DataJpaTest coverage at all —
findByEmail and the email column's
unique constraint had never been verified against a real
database. Added tests including a real proof the unique
constraint is genuinely enforced.
itemCode field CodeQL
had already flagged once (item 06) turned out to have a second,
cross-service path to an unsanitized log line:
order-service → payment-service →
restaurant-service's raw string-concatenated
rejection reason → back into order-service's own
cancelOrder. Traced end-to-end before being fixed
with the same sanitizing pattern already established for the
first instance.
customerId (and
one carried ticketId) all the way into
confirmOrder/cancelOrder without
either method ever checking them — a real asymmetry against
this repo's own established pattern of validating every field
on an inbound event. Fixed by cross-checking
customerId against the order already on record,
catching a corrupted or mismatched message on the shared topic
instead of trusting it blindly.
Future.get(), implicitly relying on Kafka's own
undocumented 120-second client default while holding a database
row lock open the whole time — a degraded broker could have
held that lock for up to 20 minutes across a full batch. Fixed
with an explicit, bounded timeout.
PaymentStatus.FAILED was declared but dead code —
processPaymentSaga unconditionally approved every
payment, so the saga's own decline/compensation path was
unreachable. Added a real, deterministic decline rule (a
configurable authorization limit) and a guard so a declined
payment can never be incorrectly refunded, cascading a matching
fix into restaurant-service ahead of that module's
own audit turn.
handleOrderCompensation refunded a payment using
only the order ID off the inbound event, never cross-checking
the customer ID against the payment it was refunding — unlike
order-service's identical compensation path, which already
guards against exactly this class of corrupted/mismatched Kafka
message. Added the same cross-check.
InventoryItem row outside of tests — no seed
script, no replenishment endpoint. Every real order would hit
ITEM_NOT_FOUND unconditionally, the module's core
function unreachable in any actual deployment. Added seed data
covering all three outcomes, and caught a genuine H2/Postgres
SQL portability bug while verifying it.
RestaurantService uses an existsByOrderId
check-then-act idempotency guard against duplicate
PaymentProcessedEvent delivery — a real occurrence
under Kafka's default at-least-once semantics, not theoretical —
but the order_id column itself had no unique
constraint backstopping it, unlike payment-service's
identical guard. Fixed by matching that existing precedent, plus
a test proving a duplicate insert is genuinely rejected.
./gradlew test aggregateTestReport in one
invocation could intermittently fail Gradle's own safety
validation: the report task read each module's test-results
directory without any declared ordering relationship to the
test task that writes it, so a same-run schedule
could read a directory before it was ready. Reproduced reliably,
then fixed with an explicit ordering hint that doesn't force
tests to run or block the report on a failure — preserving the
task's deliberate "merge whatever results exist, pass or fail"
design.
kafkaTemplate.send(message).get()
blocked unbounded on Kafka's own undocumented 120s default while
holding a database row lock open the whole time, identical to
the bugs already fixed in order-service and
payment-service. Fixed with an explicit, bounded
timeout. Verified via deliberate revert — reverting makes the
new test fail to compile, not just fail, a genuine
compile-time proof.
DefaultErrorHandler anywhere in the module — an
uncaught exception from a malformed payload fell through to
Spring Boot's autoconfigured default (10 retries at 0ms
backoff, then silent log-and-skip). The same gap already fixed
in order-service and payment-service's
own Kafka consumers. Fixed with a bounded backoff and a
non-retryable exception list, verified via deliberate revert.
verifyAndDeductStock read an item's stock with a
plain findById, then wrote the deduction back with
no locking — two concurrent orders for the same item could both
read the same stock count, both pass the check, and both get
allocated. Reachable once the service scales past one instance,
since the Kafka message key is orderId, not
itemCode. Proved with a real
@SpringBootTest firing two genuinely concurrent
calls — [ALLOCATED, ALLOCATED] without the fix,
[ALLOCATED, INSUFFICIENT_STOCK] with it. Fixed with
a pessimistic-write row lock.
Mono.fromCallable(...) with no scheduler
override, on a comment's mistaken claim that
spring.threads.virtual.enabled kept it off
Reactor Netty's event loop — it doesn't; that setting only
retargets servlet-style executors. Verified empirically
with a real running server: the call executed on
webflux-http-nio-2, a genuine event-loop
thread, meaning every login could block one of a handful
of threads shared by the whole gateway. Fixed with
subscribeOn(Schedulers.boundedElastic()).
java-jwt source jar confirmed
JWTDecodeException embeds the raw,
attacker-controlled, base64url-decoded token segment
verbatim into its own message — so a
Bearer token whose payload decoded to
non-JSON text containing literal CR/LF could forge fake
log lines, unauthenticated, before signature
verification ever ran. Fixed by sanitizing the message
before logging, the same pattern already used for
OrderService's identical finding.
@RequestBody
record field with no character restrictions, so a JSON
body's CR/LF escapes could forge fake log lines on an
endpoint that needs no authentication to reach at all.
Fixed with the same sanitizing pattern used everywhere
else this class of bug turned up, plus a real test class
added where none existed before — the client's own logic
had zero coverage anywhere in the suite.
TimeLimiter, then re-verified from a fresh
stack restart that the very first cold call succeeds.
SUCCESS still saw it flicker back through
PENDING first. Root cause:
Sinks.many().multicast().onBackpressureBuffer()
queues any emission made while nobody's subscribed and replays
the whole backlog to the next arrival, regardless of relevance.
Fixed by switching to directBestEffort(), which only
delivers to a subscriber connected at the exact moment of
emission.
MethodRoutePredicateFactory matches a request's
literal HTTP method with zero preflight-awareness — a route
restricted to POST or GET never
matches a real browser's OPTIONS preflight at all.
Confirmed via bytecode inspection of
AbstractHandlerMapping.getHandler(): its
CORS-processing step only runs once a route has already
matched, so an unmatched preflight falls through unrouted
regardless of how the CORS config itself is written. Fixed by
adding OPTIONS to every guarded route's method
predicate — safe, since a matched preflight is still
short-circuited before any route filter runs.
/auth/login kept getting
rejected with 403 even though the gateway's CORS config looked
correct on paper — the identical origin worked fine on a
gateway-routed path. Diagnosed empirically with a temporary
diagnostic WebFilter before writing any fix:
Spring's own @RestController dispatch
(AuthenticationController) claims
/auth/login before the gateway's own routing ever
sees it, and the CORS config was only ever wired into the
latter — so the path actually being served had no CORS
configuration behind it at all, invisible from reading the YAML
alone. Fixed with a HandlerMapping-agnostic
CorsWebFilter that applies uniformly regardless of
which mapping serves a request. A real Dockerfile inefficiency
found in the same debugging pass — no BuildKit Gradle cache
mount, costing 10+ minutes per rebuild — was fixed in the same
change.
Real Postgres, real Kafka, real partition assignments — not assumed
Every Kafka-wired service was booted for real against Docker Compose's
Postgres and Kafka containers, not just unit-tested in isolation: a
genuine HikariPool→PgConnection connection,
and each service's consumer groups actually joining and getting real
partition assignments against the live broker — the thing every prior
test run under a mocked or absent broker couldn't prove. Every service
now runs as its own container too, not just its infra dependencies —
the full 7-container stack (5 services + Postgres + Kafka) starts with
one docker compose up, verified end to end including the
saga's compensation path, not just the happy path.
Repository tests run against embedded H2 with real Hibernate DDL, not mocks; Kafka/Postgres wiring is separately verified against the actual Docker containers before being called done.
A consolidated test report, regenerated and committed back to main by CI after
every real test run — checkable against the live repository, not
a static screenshot.
Where the interesting decisions actually happened
GitHub Pages was briefly switched to an Actions-based deployment to
try matching this portfolio's own live-rendered-site pattern, then
deliberately reverted back to a simpler committed-static-file
approach once it became clear the fancier setup would replace the
README-rendered repo homepage for no real benefit at this stage —
matching a pattern isn't a reason to add complexity a project doesn't
need yet. Separately, building out payment-service before
restaurant-service wasn't the obvious reading of the
original four-service plan — it came from actually checking which
service consumes which event, not from the plan's stated order.
restaurant-service depends on payment-service's
output, not the reverse, so building it first would have meant
building against nothing. Both calls, and the reasoning behind them,
are recorded in
todo.md
as they happened, not reconstructed afterward.