Case study · 2026 · FlyRank AI Capstone · Back-End AI Engineering Intern
Usage Metering & Billing Engine
Billing code fails in one of two directions: it overcharges a real customer, or it gives service away for free. Both are caused by the same thing: a retry, a replayed webhook, or an off-by-one at the quota boundary. This service treats those three as the actual product and makes the database, not application logic, the thing that guarantees correctness.
The problem
Every SaaS has to answer three questions for every tenant, every month: how much have they used, what should they pay, and have they hit their plan limit. They look like three separate features. They are really one correctness problem, because every one of them breaks under the same conditions: a client retrying a request it already sent, Stripe redelivering an event it already delivered, or a request landing exactly on the quota line.
A bug in any of those either bills a real customer for work they did not do, or hands out service nobody paid for. So the design started from the failure modes rather than the endpoints.
Exactly-once metering, enforced by the database
Every billable call carries an Idempotency-Key header, and the usage_events table holds a UNIQUE constraint on (tenant_id, idempotency_key). That constraint (not an in-memory check, not application logic) is the exactly-once guarantee, which means it still holds when two identical retries race each other concurrently.
The lookup path returns the original response for a duplicate key, with no second quota check and no new row. The race path matters more: if two concurrent retries both get past the lookup, one loses on the unique constraint at insert time. That violation is caught and treated as a replay rather than an error, so the loser of the race still receives a correct and consistent answer instead of a 500.
Stripe webhook dedup works on the same principle: a UNIQUE constraint on stripe_event_id in a webhook_events table. Verified under a genuine replay, not a simulated one: a stripe events resend delivered the same real event twice through the live CLI forward, and it was processed once and ignored the second time.
Money math in integers
AI token prices are fractions of a cent, which is exactly where floating point quietly corrupts a ledger. Prices are pinned as integer micro-cents per token (1 cent = 1,000,000 micro-cents), so fractional-cent-per-token rates stay exact integers all the way through the rollup with no rounding drift.
Cached input tokens bill at 5 micro-cents against 20 for fresh input, and reasoning tokens bill at the output rate rather than becoming a fourth pricing category, a deliberate choice to keep the pricing table small enough to reason about.
Drawing the quota boundary once
The boundary rule is stated exactly once in the design and never restated: a request that brings usage to precisely the limit is allowed; the next one is rejected. On a 1,000-call plan, the 1,000th call passes and the 1,001st is refused.
Writing it down once, rather than re-deriving it at each call site, is what makes it testable. Three of the 20 tests exist purely to pin that boundary in place.
Scope discipline
Overage billing, proration, and invoicing were cut on purpose. Usage past quota is rejected, not billed extra. That keeps the money math to one thing, metered token pricing, done correctly, rather than three things done approximately.
The same honesty applies to what is missing. There is no reconciliation job to catch a webhook Stripe tried to deliver and failed to, and the background alert worker logs structured text where a real deployment would want JSON lines going to an aggregator. Both are written down in the repo as known gaps rather than quietly omitted.
Outcome
Accepted by the lead track mentor as the culminating project of the FlyRank AI Backend Engineering internship. 20 passing tests cover idempotent metering, the exact quota boundary, eight token-pricing cases, and webhook signature verification including forged-signature rejection.
Beyond the test suite, every requirement was run once end to end against a real Stripe test-mode account: a real Checkout session paid with a test card, a real webhook flipping the tenant from Free to Pro, a forged signature header rejected by the live server, and a genuine Stripe event replay processed exactly once.
The track behind it
Six assignments preceded the capstone, each one a standalone repo. They move from a plain CRUD API to durable AI workflows, and the reasoning in each README is the part worth reading.