Single Redis to Redis Cluster. 38 GB working set. What we planned (3 weeks) vs. what happened (9 weeks). What we discovered: 200+ multi-key transactions with runtime-variable keys, 8 Lua scripts spanning slots, rate limiter with slot-migration ambiguity, a 3-year-old leaderboard bug silently corrupting weekly totals. Migration was an application audit, not a database migration.
Practitioner posts on shipping with Claude Code
Opinion, setup guides, patterns, cost analysis, deep dives. Written by engineers doing the work, not observers commenting on it. New posts weekly.
Latest
Freshest post from the last week.
SOC2 audit revealed 340 credentials, 200+ without owners, 60+ without rotation. Anomalies: shared DB password used by 20 people, ex-employee's AWS key still valid 8 months after departure, encryption key on a public S3 bucket, 'dev' secret that was actually prod. What we retrofitted: catalog + owner assignment, per-class rotation cadence, tag enforcement via hook, break-glass separation, secret sc
All posts
Filter by category or author. Search across titles and summaries.
Boolean columns hide payment bugs; state machines surface them. Explicit states + explicit transitions + rejected illegal transitions. The pattern we settled into after a trial-refund-upgrade bug. Real anomalies caught (dunning after cancellation, out-of-order webhooks); reconciliation with Stripe; test coverage tractable; new engineer onboarding faster; business metrics fall out of the model.
Friday rolling deploy at 3 PM UTC. 120k WebSocket connections. Client-side reconnect logic without backoff. What we expected (2s blip per user), what actually happened (6-min compounding retry storm), what we shipped in the moment (slower rolling deploy), what we designed after (client backoff+jitter, server reconnection budget, ALB throttling, health-gated deploy advancement, auth check caching),
We had 843 feature flags. This is the opinion piece: sunset is engineering, not cleanup. Every flag has cognitive + testing + latency + reasoning cost. Flag typology matters (release, experiment, ops, permission, circuit breaker — not all should die). The 2-month cleanup approach (categorize, identify zombies, remove, review). The ongoing discipline that prevents regrowth. Two near-incidents
Partner's misconfigured retry logic sent us 340k requests in 4 minutes. We didn't have gateway rate limits. What we shipped in the moment (Cloudflare block for 3 min), what we designed after (Kong per-partner tiers + sliding window counter + shadow-mode rollout + per-endpoint overrides + rate limit headers + partner status page), what the same architecture caught 3 months later (competitor scrapin
6 months post-launch of WebAuthn passkeys. Adoption: 40%. Support tickets: elevated first month then baseline. Phishing incidents: zero among enrolled. Surprising positive: password reset support tickets dropped 78% for enrolled cohort. What we shipped wrong first (missing explanation, missing multi-device UX), what we changed, what we'd do differently on next rollout.
Actual incident: 22-min partner outage → 380k webhook backlog → retry storm → DB primary at 98% CPU → 22 min API down + 3 hours degraded. What we thought first (all wrong), what actually happened (retry compounding across partners), how we mitigated (per-partner concurrency limits), what we changed structurally, what we almost broke doing the fix.
A single incident 18 months ago changed our post-mortem practice. Not the response (fine) but what came after. Wrote the wrong post-mortem first (blamed people, vague action items, closed ticket). Same incident happened again 63 days later. What we changed: blameless framing, designated author, cross-team publishing, quarterly meta-review. What shifted (repeat incidents dropped, action item comple
Five-week caching journey: Postgres primary at 78% CPU. Attempted broad caching (didn't work), query optimization (partial), read replica (expensive). Landed on tier-specific caching of top 12 queries. Load 78% → 31%. What we monitored, what we almost broke, cost impact ($180/mo vs. $1700/mo alternatives).
9-month migration from Datadog-only to OpenTelemetry with multi-backend split (Datadog + Grafana LGTM). What we saved ($16k/mo), what we broke, what we'd do differently. Case study in vendor optionality, not Datadog critique.
Six-month journey from synchronous checkout to event-driven with Kafka + Temporal + OpenTelemetry. What worked (latency, reliability, deploys); what took 2x longer (idempotency, sagas, schema evolution); what we'd do differently. Practical lessons, not advocacy.
Design system as team contract
Reframing design system from library (optional shared components) to contract (rules everyone agreed on). Explicit ownership, deprecation cycles, structured feedback, automated enforcement. What Claude Code made cheaper; what still requires humans.
The monorepo tools we tried before landing on Turborepo
14 months, four tools: Lerna, Nx, pnpm-only, Turborepo. Why each didn't fit until Turborepo did. What we monitor now (cache hit rate 85%); what would change our decision. Trade-off framework, not tool advocacy.
Our internationalization strategy after supporting 8 locales
Rebuilt from scratch after adding locales 4, 5, 6 broke things. Seven pieces: ICU MessageFormat, CSS logical properties, context per key, automated extraction, Intl APIs, pseudo-locale testing, orphan cleanup. What Claude Code changed; what still doesn't work.
Accessibility is engineering, not compliance
Reframing accessibility as engineering discipline, not compliance sprint. Baseline + block regressions; component library as contract; automated + keyboard + screen reader; what Claude Code changed. From 20 months of practice.
Why we picked GitLab over GitHub for internal tooling
Not a hot take; specific reasoning. GitLab for internal (integrated CI/registry/K8s; self-hosted; small ops); GitHub for open source (network effect; community). Honest section on what GitHub does better and what would flip us back.
The release engineering discipline we settled into
18 months of iteration on release engineering. Seven pieces: conventional commits, automated changelogs, semver via semantic-release, narrative release notes, weekly release trains, canary rollout, quarterly rollback rehearsal. What we do; what we explicitly don't.
Linear + Claude Code: our ticket-to-PR workflow
Concrete setup for Linear MCP + Claude Code, three workflows that changed daily (ticket start, commit+PR+status, review context), what we explicitly don't do, metrics after 6 weeks. Applies to Jira/Asana with different MCP.
How we cut p99 latency 40% with Claude Code
5-week investigation: p99 latency drifted from 380ms to 520ms; systematic scan with performance-analyzer + Prometheus MCP brought it to 310ms. Breakdown of contributions: N+1 fix, tax parallelization, caching, Redis upgrade. Investigation discipline that worked.
The security audit playbook we actually run
Six pieces: diff-level review, weekly scans, quarterly deep audit, yearly pen test, incident-driven audit, onboarding audit for major deps. Boring. Sustainable. Actually run. Also: the things we explicitly don't do and why.
What Prometheus MCP unlocked for our on-call
First on-call rotation after installing Prometheus MCP: 3 of 5 pages resolved faster. Not because Claude did anything I couldn't; because it correlated in parallel while I was still reading. What changed, with specific cases.
The mock discipline that made our tests fast again
4 minutes to 45 seconds. Fixed three mock discipline issues: mocking too deep, missing cleanup, over-specification. 5x faster; 4x less flaky; refactoring velocity up. The specific fixes with numbers.
When we write RFCs (and when we don't)
The threshold that worked for our 10-person team: two-of-four criteria (documented dissent might matter, affects people not in the meeting, reversal cost is high, decision will outlast current staff). Six RFCs in twelve months; here's the list.
Why we scan for PII before every commit
A near-miss with 40 real customer emails in a test fixture prompted us to install content-level PII scanning. Four months in: 47 detections, mostly caught legitimate mistakes. Prevention vs. incident response calibration.
Hunting flaky tests: the pattern that finally worked
47 flaky tests → 3 in six months. Not retries, not deletion, not stoicism. Root-cause classification with six patterns; fix at root cause; specific results. CI success rate 82% → 97%; retries 2.3 → 0.2 per PR.
How we do database migrations safely with Claude Code
The three-step discipline (plan, generate, test) our team evolved for safe migrations with AI assistance. Multi-phase decomposition for breaking changes; specific safeguards; what Claude adds and what it doesn't.
The approval-hook pattern: our guardrail against expensive mistakes
After Claude Code helped drop a production database, we implemented a dangerous-command approval hook. This is the specific pattern, why the friction is worth it, and what we've learned about AI velocity + human judgment.
Design-to-code with Figma MCP: our handoff workflow
Four months into using Figma MCP for design-to-code handoff. Specific workflows for component implementation, token sync, icon export. What worked, what didn't, and the 15% time savings we measured.
Auto-generating our public API docs from Claude Code sessions
The specific pipeline we set up to fix embarrassingly-drifted API docs. /api-docs slash command + scheduled regeneration + drift-detection hook. Docs accuracy from 60% to 95%; developer effort from 'remember' to 'nothing.'
AWS MCP + terraform-planner: our infra investigation workflow
How we combine AWS MCP for cloud investigation with terraform-planner for IaC changes. Five scenarios covering rightsizing, security hardening, drift reconciliation, incident investigation, cost optimization.
When Claude Code should NOT touch the codebase
Six categories of work where AI assistance is worse than no assistance. Not because Claude can't handle them, but because the failure mode is worse than the alternative. Includes framework for drawing your team's own lines.
The spec-to-code loop: /spec → /adr → shipping
A deep dive into the workflow we've evolved for taking substantial changes from 'we should do this' to 'it's shipped.' Two Claude Code commands do most of the work; the discipline around them is where the value is.
Setting up Stripe MCP for a production SaaS
How we set up Stripe MCP for our team — restricted keys, test-mode default, allowlist/denylist config, per-developer credentials, rules for live-mode access. The setup we've run in production for three months.
Why we banned :latest tags on production images
A short opinion post about a specific policy that saved us from a specific class of production incident. Three-layer enforcement (CI + Claude Code hook + K8s admission webhook) makes :latest-tagged production deploys impossible. Boring, effective, one afternoon of setup.
How we generate release notes in 5 minutes with Claude
The complete setup for turning weekly release notes from a 45-minute manual chore into a 5-minute review-and-commit workflow. changelog-writer subagent + wrapper slash command + standardized CHANGELOG.md format. Three months in with the pattern; adoption rate up; support tickets down.
The architect-reviewer pattern: adding design-level review to your PR flow
How to add design-level review to your PR process without depending on senior architects catching things by chance. The architect-reviewer subagent runs Opus-based architectural review on significant PRs; complements code-reviewer and security-reviewer with a design lens. Real findings + severity discipline + when to invoke.
Why we route Claude to a dedicated Snowflake warehouse (and how)
The four-layer cost setup we use to keep Claude-mediated Snowflake queries under $12/month for a 10-person team. Dedicated small warehouse + read-only role + MCP statement allow-list + query timeout. Structural cost control that makes surprise bills impossible.
The Database Trifecta: MCP + Subagent + Migration Planner
A three-component pattern for using Claude Code with any database. The MCP gives Claude your schema, the subagent gives it DB-specific expertise, the migration planner gives it safety. Uses Postgres as exemplar; generalizes to MySQL, Mongo, and beyond.
The 3 Hooks Every Claude Code Repo Should Have
After installing Claude Code on 200+ repos, three hooks stand out as non-negotiable. Not the ones you'd guess; not the ones the docs feature. The three that catch the mistakes that actually cost teams money, ranked by ROI with the case for each.
How to Configure Claude Code for a Next.js 15 Team in 20 Minutes
The exact CLAUDE.md, slash commands, hooks, subagents, and MCPs to install for a Next.js 15 team. Copy-pasteable configs, timed to 20 minutes end-to-end. No fluff, no exploratory questions — just the working setup.
Slash Commands vs Subagents in Claude Code — When to Use Which
Decision framework for slash commands vs subagents in Claude Code. Key differences in context isolation, invocation model, cost profile, and output shape. Concrete patterns with examples.
The .env Problem — How Claude Code Teams Handle Secrets
Three concrete patterns for managing secrets in Claude Code teams: env-file-protect + .env.example, injected-at-runtime via 1Password/Vault, and project-scoped MCPs. Tradeoffs, failure modes, and the pattern we recommend.
The Real Cost of Claude Code at Team Scale — What Actually Drives Spend
Breakdown of Claude Code costs at team scale. Where the money actually goes (subagents, context, MCP loops), what teams underestimate, and concrete fixes with cost-control hooks. Real numbers from a 10-person team.
No posts match your filters
Try clearing filters or searching for a different term.
Get new posts in your inbox
Tuesday mornings. Practitioner-only content. Unsubscribe with one click. No AI hype, no vendor pitches.