Portfolio Manifest — Last updated 2026-09-09

Terrence Daniels

Backend & Distributed Systems Engineer

An original implementation of the Distributed Saga pattern across six independent Spring Boot / gRPC microservices — a JWT-guarded API gateway in front of order placement, payment settlement, and restaurant fulfillment, coordinated through Kafka with compensating transactions instead of a shared database transaction. Built and verified one service at a time against real Postgres and Kafka infrastructure — not mocked past the point where mocking stops proving anything.

56phases shipped
181/181tests passing
6services complete
3saga chain hops
31real bugs found & fixed
2real infra, not mocked

Why this project

A whiteboard topic, built for real

Saga orchestration is a common interview-whiteboard topic and an uncommon thing to actually build end to end: compensating transactions, event ordering, idempotency, and the failure modes that only show up once services genuinely run independently. This isn't a fork of any existing project — the module boundaries and general shape of the problem (order → payment → fulfillment, with compensation on failure) are common territory for this class of system, but the code, design decisions, and tradeoffs recorded here are this repo's own, built one service at a time with the reasoning written down as it happened.

Spring Boot gRPC Kafka PostgreSQL Gradle (Kotlin DSL) JUnit 5 Docker Compose GitHub Actions CodeQL

Six services, one saga chain

Outbox pattern, compensating transactions, wired end-to-end behind a real security perimeter

api-gateway-service fronts the chain — JWT-guarded routing and a Resilience4j circuit breaker sit in front of everything below. order-service creates an order and stages an outbox record in the same transaction; a poller ships it to Kafka. payment-service consumes it, settles payment, and publishes its own event. restaurant-service consumes that, allocates inventory, and makes the saga's real decision — approve or reject — publishing back to both order-service and payment-service so a rejection compensates the payment and cancels the order. No shared database transaction anywhere in the chain.

Deliberate deviations, not defaults

spring-kafka directly instead of a Spring Cloud Stream binder — this project only ever targets Kafka, so the broker-swappability abstraction wasn't worth the config overhead. UUID ids and BigDecimal money instead of raw strings and doubles — floating point can't exactly represent decimal fractions, a real correctness concern once money's involved. Reasoning for both recorded in docs/architecture.md.

Idempotency guards, not TODOs

Every state-changing method follows the same cheap-check-before-load pattern: an existsBy… query first, then a status re-check after load, so a duplicate Kafka delivery — the normal case for at-least-once delivery, not an edge case — can't double- apply an order confirmation, a payment, or a compensation.

Real bugs, not just features

Found via actually running the thing, not reading the source twice

Every entry below was caught by a real, failing run — a test, a CI job, or a live boot against real infrastructure — not by static reading, and every fix was re-verified with another real run before being called done.

01
user-service — protobuf-java silently downgraded
spring-grpc's BOM was silently pulling protobuf-java below what user-contract's generated code needs, on user-service's actual runtime classpath — not just tests — since the module's first phase. Undetected until a test finally constructed a generated message at runtime. Fixed with an explicit version pin.
high · production
02
order-service — Spring Boot 3.4.1 can't boot under JDK 25
The module's first test (contextLoads) failed outright — Spring Framework's own bundled ASM can't parse JDK 25 class files during component scanning at all, deeper than an earlier Gradle-plugin-only version of the same issue. Checking user-service showed it had the identical latent risk, never caught because no test there had ever booted a real ApplicationContext. Fixed by bumping both modules to Boot 3.5.16.
high · infra
03
Consolidated test report — two report bugs
The custom root task merging every module's JUnit XML into one HTML file had two real bugs: a failing test silently blocked the report from generating at all, and — separately — the task had no declared inputs, so Gradle kept serving a stale report instead of regenerating it. Both fixed; the report is now a genuine CI artifact on every run.
medium · tooling
04
CI — gradlew committed without its executable bit
The first real CI run failed outright (Permission denied, exit 126) — gradlew had been committed from Windows without its executable bit, so git stored mode 100644 instead of 100755. Fixed with git update-index --chmod=+x, confirmed on a second real run.
medium · CI
05
api-gateway-service — a ported test that could never pass
Porting JwtPerimeterGuardGatewayFilterFactoryTest and actually running it — never assumed — showed two of its three cases asserting an HTTP response status the filter itself never sets; that translation only happens through GlobalExceptionHandler, a layer absent from an isolated filter unit test. Failed immediately with onError instead of the expected onComplete. Rewritten to assert the filter's real, provable contract instead.
medium · test correctness
06
order-service — log injection (CWE-117)
GitHub CodeQL flagged OrderController logging a whole request record whose itemCode field is free-form client input with no character restriction — \r/\n bytes could forge fake log lines. Fixed by logging fields individually and stripping control characters from the one field that actually needed it.
medium · security
07
user-service — BCryptPasswordEncoder.encode() never exercised
No registration RPC exists in this repo, and every test mocked PasswordEncoder with a literal placeholder hash — so the real bcrypt encode() path had never run once, in production code or tests. Added a real encode/matches round-trip test.
low · test coverage
08
user-service — login enumeration (CWE-203)
Login returned three distinguishable gRPC statuses depending on unknown-email vs. wrong-password vs. inactive-account, letting a caller enumerate registered emails — and their active status — without ever guessing a password. Collapsed to one generic UNAUTHENTICATED response after weighing the UX tradeoff.
medium · security
09
user-service — login timing side-channel (CWE-208)
Even after closing the status-code leak above, unknown-email and inactive-account logins skipped the deliberately-slow BCrypt comparison entirely, responding measurably faster than a real login attempt. Restructured so every path pays the identical real BCrypt cost, verified by reverting the fix and confirming new tests fail against the old code first.
medium · security
10
user-service — JWTVerifier rebuilt on every call
The sibling code in api-gateway-service already caches its JWTVerifier once at construction for performance; this file rebuilt one on every single verifyToken() call. Matched the established pattern — zero behavior change, existing tests passed unchanged as proof.
low · maintainability
11
user-service — zero repository test coverage
Unlike every other module's repositories, user-service had no @DataJpaTest coverage at all — findByEmail and the email column's unique constraint had never been verified against a real database. Added tests including a real proof the unique constraint is genuinely enforced.
low · test coverage
12
order-service — Kafka poison-pill risk, no error handler configured
Neither Kafka listener caught anything — an event referencing an unrecognized order ID or a malformed payload fell through to Spring Boot's autoconfigured default with nothing in the code announcing that. Verifying that default for real turned up a second bug: an earlier write-up had already mischaracterized it as retrying indefinitely, when the actual behavior is 10 rapid retries then a silent drop. Fixed with an explicit, bounded error handler — and the inaccurate claim corrected in the same PR.
medium · reliability
13
order-service — a second log injection, three services away from the first
The same client-controlled itemCode field CodeQL had already flagged once (item 06) turned out to have a second, cross-service path to an unsanitized log line: order-service → payment-service → restaurant-service's raw string-concatenated rejection reason → back into order-service's own cancelOrder. Traced end-to-end before being fixed with the same sanitizing pattern already established for the first instance.
medium · security
14
order-service — customerId and ticketId deserialized, never read
Two inbound Kafka events carried customerId (and one carried ticketId) all the way into confirmOrder/cancelOrder without either method ever checking them — a real asymmetry against this repo's own established pattern of validating every field on an inbound event. Fixed by cross-checking customerId against the order already on record, catching a corrupted or mismatched message on the shared topic instead of trusting it blindly.
medium · reliability
15
order-service — outbox publisher blocked on Kafka with no timeout
The transactional outbox's Kafka send blocked on an unbounded Future.get(), implicitly relying on Kafka's own undocumented 120-second client default while holding a database row lock open the whole time — a degraded broker could have held that lock for up to 20 minutes across a full batch. Fixed with an explicit, bounded timeout.
low · reliability
16
payment-service — a payment could never actually fail
PaymentStatus.FAILED was declared but dead code — processPaymentSaga unconditionally approved every payment, so the saga's own decline/compensation path was unreachable. Added a real, deterministic decline rule (a configurable authorization limit) and a guard so a declined payment can never be incorrectly refunded, cascading a matching fix into restaurant-service ahead of that module's own audit turn.
medium · reliability
17
payment-service — refund compensation never checked whose order it was refunding
handleOrderCompensation refunded a payment using only the order ID off the inbound event, never cross-checking the customer ID against the payment it was refunding — unlike order-service's identical compensation path, which already guards against exactly this class of corrupted/mismatched Kafka message. Added the same cross-check.
low · reliability
18
restaurant-service — no way to ever seed inventory in a real run
No path anywhere in the repo could create an InventoryItem row outside of tests — no seed script, no replenishment endpoint. Every real order would hit ITEM_NOT_FOUND unconditionally, the module's core function unreachable in any actual deployment. Added seed data covering all three outcomes, and caught a genuine H2/Postgres SQL portability bug while verifying it.
medium · reliability
19
restaurant-service — duplicate Kafka delivery could create two tickets for one order
RestaurantService uses an existsByOrderId check-then-act idempotency guard against duplicate PaymentProcessedEvent delivery — a real occurrence under Kafka's default at-least-once semantics, not theoretical — but the order_id column itself had no unique constraint backstopping it, unlike payment-service's identical guard. Fixed by matching that existing precedent, plus a test proving a duplicate insert is genuinely rejected.
medium · reliability
20
build tooling — the consolidated test report could fail to generate
Running ./gradlew test aggregateTestReport in one invocation could intermittently fail Gradle's own safety validation: the report task read each module's test-results directory without any declared ordering relationship to the test task that writes it, so a same-run schedule could read a directory before it was ready. Reproduced reliably, then fixed with an explicit ordering hint that doesn't force tests to run or block the report on a failure — preserving the task's deliberate "merge whatever results exist, pass or fail" design.
low · tooling
21
order-service — a 10-day-old CodeQL alert had never actually been resolved
A log-injection alert flagged since PR #55 stayed open through two follow-up fix attempts, each leaving a code comment admitting it was "still flagged" without ever closing the loop. Investigated properly this time: two more fix attempts, each verified against a real CodeQL run, each cleared the prior alert only to reveal a new one at the next line — confirming the value was already being sanitized correctly the whole time. Researched CodeQL's own sanitizer source to confirm it, then formally dismissed the alert with a documented reason instead of leaving another silent "still flagged" comment behind. The branch split and shared sanitizer helper written along the way are kept as real clarity improvements.
low · security
22
restaurant-service — outbox publisher blocked on Kafka with no timeout
The third occurrence of the same gap: kafkaTemplate.send(message).get() blocked unbounded on Kafka's own undocumented 120s default while holding a database row lock open the whole time, identical to the bugs already fixed in order-service and payment-service. Fixed with an explicit, bounded timeout. Verified via deliberate revert — reverting makes the new test fail to compile, not just fail, a genuine compile-time proof.
medium · reliability
23
restaurant-service — no error handling on the Kafka consumer
No DefaultErrorHandler anywhere in the module — an uncaught exception from a malformed payload fell through to Spring Boot's autoconfigured default (10 retries at 0ms backoff, then silent log-and-skip). The same gap already fixed in order-service and payment-service's own Kafka consumers. Fixed with a bounded backoff and a non-retryable exception list, verified via deliberate revert.
medium · reliability
24
restaurant-service — unlocked stock-deduction race
verifyAndDeductStock read an item's stock with a plain findById, then wrote the deduction back with no locking — two concurrent orders for the same item could both read the same stock count, both pass the check, and both get allocated. Reachable once the service scales past one instance, since the Kafka message key is orderId, not itemCode. Proved with a real @SpringBootTest firing two genuinely concurrent calls — [ALLOCATED, ALLOCATED] without the fix, [ALLOCATED, INSUFFICIENT_STOCK] with it. Fixed with a pessimistic-write row lock.
high · concurrency
25
api-gateway-service — blocking login call on the Netty event loop
The login endpoint's blocking gRPC call was wrapped in Mono.fromCallable(...) with no scheduler override, on a comment's mistaken claim that spring.threads.virtual.enabled kept it off Reactor Netty's event loop — it doesn't; that setting only retargets servlet-style executors. Verified empirically with a real running server: the call executed on webflux-http-nio-2, a genuine event-loop thread, meaning every login could block one of a handful of threads shared by the whole gateway. Fixed with subscribeOn(Schedulers.boundedElastic()).
medium · reliability
26
api-gateway-service — CWE-117 log injection via a malformed JWT's decoded payload
The JWT filter's catch-all handler logged a failed token's exception message raw. Extracting auth0's java-jwt source jar confirmed JWTDecodeException embeds the raw, attacker-controlled, base64url-decoded token segment verbatim into its own message — so a Bearer token whose payload decoded to non-JSON text containing literal CR/LF could forge fake log lines, unauthenticated, before signature verification ever ran. Fixed by sanitizing the message before logging, the same pattern already used for OrderService's identical finding.
high · security
27
api-gateway-service — a second CWE-117 log injection, this time on the login endpoint itself
The gRPC login client logged the submitted email raw, twice, on every attempt — a @RequestBody record field with no character restrictions, so a JSON body's CR/LF escapes could forge fake log lines on an endpoint that needs no authentication to reach at all. Fixed with the same sanitizing pattern used everywhere else this class of bug turned up, plus a real test class added where none existed before — the client's own logic had zero coverage anywhere in the suite.
high · security
28
api-gateway-service — gateway circuit breaker had no explicit call timeout, silently defaulting to 1 second
Found while bringing the whole stack up in real Docker containers for the first time: the very first order request — a cold cross-container HTTP call, DNS resolution and connection-pool warmup included — took 1.29 seconds and got a false-positive circuit-open fallback under Resilience4j's undocumented library default. Fixed by configuring an explicit 5-second TimeLimiter, then re-verified from a fresh stack restart that the very first cold call succeeds.
medium · reliability
29
order-service — SSE stream replayed stale buffered updates to a late-connecting client
Real end-to-end verification (not caught by any unit test — every test subscribes before the emission it's checking) showed a client connecting after an order had already reached SUCCESS still saw it flicker back through PENDING first. Root cause: Sinks.many().multicast().onBackpressureBuffer() queues any emission made while nobody's subscribed and replays the whole backlog to the next arrival, regardless of relevance. Fixed by switching to directBestEffort(), which only delivers to a subscriber connected at the exact moment of emission.
medium · concurrency
30
api-gateway-service — a CORS preflight route Gateway itself could never match
Spring Cloud Gateway Server WebFlux's own MethodRoutePredicateFactory matches a request's literal HTTP method with zero preflight-awareness — a route restricted to POST or GET never matches a real browser's OPTIONS preflight at all. Confirmed via bytecode inspection of AbstractHandlerMapping.getHandler(): its CORS-processing step only runs once a route has already matched, so an unmatched preflight falls through unrouted regardless of how the CORS config itself is written. Fixed by adding OPTIONS to every guarded route's method predicate — safe, since a matched preflight is still short-circuited before any route filter runs.
medium · reliability
31
api-gateway-service — the login endpoint's CORS bug was a HandlerMapping race, not a config error
A real preflight to /auth/login kept getting rejected with 403 even though the gateway's CORS config looked correct on paper — the identical origin worked fine on a gateway-routed path. Diagnosed empirically with a temporary diagnostic WebFilter before writing any fix: Spring's own @RestController dispatch (AuthenticationController) claims /auth/login before the gateway's own routing ever sees it, and the CORS config was only ever wired into the latter — so the path actually being served had no CORS configuration behind it at all, invisible from reading the YAML alone. Fixed with a HandlerMapping-agnostic CorsWebFilter that applies uniformly regardless of which mapping serves a request. A real Dockerfile inefficiency found in the same debugging pass — no BuildKit Gradle cache mount, costing 10+ minutes per rebuild — was fixed in the same change.
high · reliability

Verifying it actually works

Real Postgres, real Kafka, real partition assignments — not assumed

Every Kafka-wired service was booted for real against Docker Compose's Postgres and Kafka containers, not just unit-tested in isolation: a genuine HikariPool→PgConnection connection, and each service's consumer groups actually joining and getting real partition assignments against the live broker — the thing every prior test run under a mocked or absent broker couldn't prove. Every service now runs as its own container too, not just its infra dependencies — the full 7-container stack (5 services + Postgres + Kafka) starts with one docker compose up, verified end to end including the saga's compensation path, not just the happy path.

181/181 tests, real infra behind them

Repository tests run against embedded H2 with real Hibernate DDL, not mocks; Kafka/Postgres wiring is separately verified against the actual Docker containers before being called done.

Live, committed test report

A consolidated test report, regenerated and committed back to main by CI after every real test run — checkable against the live repository, not a static screenshot.

Judgment calls

Where the interesting decisions actually happened

GitHub Pages was briefly switched to an Actions-based deployment to try matching this portfolio's own live-rendered-site pattern, then deliberately reverted back to a simpler committed-static-file approach once it became clear the fancier setup would replace the README-rendered repo homepage for no real benefit at this stage — matching a pattern isn't a reason to add complexity a project doesn't need yet. Separately, building out payment-service before restaurant-service wasn't the obvious reading of the original four-service plan — it came from actually checking which service consumes which event, not from the plan's stated order. restaurant-service depends on payment-service's output, not the reverse, so building it first would have meant building against nothing. Both calls, and the reasoning behind them, are recorded in todo.md as they happened, not reconstructed afterward.