# Mohammad Abu Mattar, Full Content Bundle
> Full text of the professional core pages and all case studies, then a link index to every content section. For the exhaustive list of Markdown twins, see /llms-sitemap.txt.
---
# Mohammad Abu Mattar
I'm Mohammad Abu Mattar, an AWS-certified DevOps engineer. Multi-account AWS for fintech: 15+ accounts, 20+ production microservices, 96% security posture.
## About
I look after 15+ AWS accounts and 20+ production microservices for fintech, under PCI-DSS. Day to day that means Terraform, GitOps with ArgoCD, and the compliance work most people would rather skip: IAM boundaries, evidence that survives an audit, controls that hold when someone is in a hurry. Production security posture currently sits at 96%.
## Experience
### Cloud & DevOps Manager, Motory · Full-time, Mar 2026 - Present, Al Hamra, Jeddah · Amman, Jordan
I run the cloud strategy and platform engineering for Motory, a large automotive marketplace. I manage a multi-cloud environment across AWS, Huawei Cloud, and Hetzner, staying about 70% hands-on with the architecture and code. I lead a DevOps team responsible for the reliability, scalability, and security of our production microservices serving 1M+ active users. I also own our CI/CD practices, cost optimization, and compliance work in a high-availability, multi-region setup.
- **Infrastructure as Code:** Moved everything from manual console clicks to 100% Terraform. I set up GitOps using ArgoCD and Helm to keep our deployments consistent and auditable.
- **Modernization:** Led the shift from a monolith to microservices for 1M+ users. I wrote a multi-runtime Helm chart that standardized how we handle scaling, probes, and networking across all services.
- **Custom Tooling:** Built a custom Kubernetes operator in Go to solve a specific provisioning problem that standard controllers couldn't handle. It plugs into the existing Helm and ArgoCD delivery flow, so the resources it manages deploy the same way everything else does.
- **CI/CD & Delivery:** Built pipelines using Jenkins and ArgoCD that improved deployment speed by 60%. This includes a private OCI registry and a Kong API gateway. I also added multi-arch builds and automated chart publishing, so releases stopped needing a person to babysit them.
- **Reliability & SRE:** We hit 99.999% uptime. I automated cross-region disaster recovery between the Middle East and China using Ansible and the Huawei SDK, and I run our observability stack (Prometheus/Grafana/Loki). On-call runbooks and regular DR drills are part of our routine.
- **Cost & Security:** Cut infra costs by 20% while tightening security to meet PCI-DSS and SOC2 standards. This involved identity-as-code via Keycloak, hardening our WAF and VPN, and moving secrets onto KMS/CSMS.
- **Team Leadership:** I lead the DevOps team and own the platform roadmap. I also run internal K8s and Docker training for our developers and handle the vendor relationships.
### Assistant Manager DevOps Engineer, Jordan Ahli Bank · Full-time, Oct 2024 - Feb 2026, Amman, Jordan
I was platform owner for the bank's AWS estate. Regulated banking means every change needs an audit trail and a rollback that works, so most of the job was making the correct path the easy one and then proving it to auditors.
- **Platform ownership:** Ran 15+ AWS accounts and 20+ production microservices, against the availability and disaster-recovery targets a bank actually gets held to.
- **Terraform & GitOps:** Standardized deployments on Terraform, with ArgoCD and Atlantis so changes landed through review instead of the console. Handled autoscaling and right-sizing for production workloads.
- **Security posture:** Got to 96% across every account using IAM Identity Center, AWS Organizations, and PCI-DSS controls backed by CloudTrail, AWS Config, and Security Hub.
- **Delivery pipelines:** Built the delivery path on AWS CodePipeline and GitHub Actions, with automated testing and security scanning before anything reached production.
- **Release engineering:** Ran weekly releases and hotfixes across dev, QA, and prod. Zero-downtime deploys, a rollback anyone on-call could run, and a real go/no-go call before shipping.
- **Production support:** Centralized monitoring on Prometheus, Grafana, and CloudWatch, plus the incident response workflow and the longer-term platform roadmap.
### DevOps Engineer, cirrusgo (AWS Partner) · Full-time, Mar 2023 - Sep 2024, Amman, Jordan
cirrusgo is an AWS partner working with businesses across the MENA region.
I was the engineer on hybrid AWS builds for 8+ enterprise clients, several of them fintech. Most of them arrived with an on-premises estate they were not going to abandon, so the work was making AWS and their existing racks behave like one platform without upsetting whichever regulator they answered to.
- **Solution architecture:** Designed the end-to-end builds (microservices, event-driven, serverless) and brought cloud costs down 25-40% through right-sizing, reserved instances, and switching off what nobody was using.
- **IaC foundations:** Set up the Terraform and Terragrunt layer and the multi-account, multi-environment structure on AWS Organizations, then wrote the documentation that went with it.
- **Container platforms:** Ran containerized workloads on ECS, Fargate, and EKS across dev, UAT, and prod, deployed through GitOps with hardened images.
- **CI/CD:** Built pipelines on GitHub Actions and AWS CodePipeline/CodeBuild with canary deploys, automated tests, manual approvals, and promotion from dev to UAT to prod.
- **Security & compliance:** Handled access control and compliance across three regions (Middle East, US East, Europe) and the on-premises side, with security scanning and monitoring running unattended.
### Software Engineer - Microservices, Nagarro · Contract, Feb 2022 - May 2022, Remote
Nagarro is a German software services company doing consulting and outsourcing work.
A short contract on the fintech side, working on Spring Boot services that moved money and the PostgreSQL layer underneath them.
- **Fintech services:** Built Spring Boot microservices for transaction processing, and stayed with them through the full SDLC rather than handing them off at merge.
- **PostgreSQL:** Owned persistence, including schema migrations and connection pooling.
- **Front end:** Built components in Angular, HTML, and CSS, and cleared production bugs and performance problems as they surfaced.
- **Documentation:** Wrote the technical docs so the next person did not have to reverse-engineer any of it.
### Full Stack Developer, Freelance · Full-time, Jan 2020 - Feb 2023, Amman, Jordan · Remote
Client work I ran end to end, alongside a full-time job for most of it. Fintech backends in Spring Boot, React front ends, and everything around the code: scoping, timelines, and the awkward conversations about what would actually ship.
- **Scoping:** Ran the conversations that turned a vague ask into something buildable, and kept clients honest with me about what they actually needed first.
- **Backend:** Designed Spring Boot microservices for fintech clients, with PostgreSQL and other databases behind them.
- **Front end:** Built responsive interfaces in React, HTML, and CSS, tested on the devices the clients' customers were really using.
- **Delivery:** Managed timelines and deliverables, and owned the client relationship directly with nobody in between.
### Teacher Assistant/Lab Supervisor, Isra University · Full-time, Dec 2018 - Dec 2019, Amman, Jordan
Isra University is a private institution in Amman offering undergraduate and postgraduate programs.
I taught C++ fundamentals and ran the lab sessions. In practice I was the person students came to when the compiler said something they could not decode.
- **Lab sessions:** Built and ran the C++ labs, and sat with students while they worked through the exercises rather than just handing them out.
- **Curriculum:** Worked with faculty to design and update the lab material.
- **Support:** Helped students and faculty with whatever programming problem was blocking them.
- **Feedback:** Tracked how students were progressing and told them plainly where they stood.
### Web Developer, Greater Amman Municipality · Apprenticeship, Dec 2017 - Feb 2018, Amman, Jordan
The Greater Amman Municipality runs the city's administration and planning.
My first job in the field: a short apprenticeship building internal web tools, mostly front-end work in HTML, CSS, and JavaScript.
- **Internal tools:** Built and maintained web applications for the administrative teams.
- **Requirements:** Sat with the team to work out what they needed before writing any of it.
- **Testing:** Tested and debugged what I shipped instead of waiting for someone to report it.
- **Training:** Supported and trained the people who had to use the tools every day.
## Education
### AWS Certified Developer - Associate, Amazon Web Services (AWS), Issued Feb 2024 · Expires Feb 2027
Covers developing, deploying, and debugging applications on AWS, plus core AWS services and the application lifecycle.
### AWS Certified Cloud Practitioner, Amazon Web Services (AWS), Issued Oct 2023 · Expires Feb 2027
Foundational AWS knowledge: core services, pricing, security, and architecture.
### AWS Academy Graduate - AWS Academy Cloud Foundations, Amazon Web Services (AWS), Issued Nov 2022
Intro-level AWS coursework covering compute, networking, databases, and storage.
### Advanced Software Development, Code Fellows, Issued Sep 2021
Full-stack development, testing, and advanced programming concepts, taught through agile team projects.
### Bachelor of Technology - BTech, Software Engineering, Isra University, Oct 2014 - June 2018
Studied core software engineering principles, data structures, algorithms, and system design. Gained hands-on experience in programming, databases, and software development life cycle (SDLC).
Achievements: First place winner in individual project at annual Technology Day contest. Contributed to establishment and growth of Faculty of Information Technology Club.
## Projects
### QuenchWorks, 0-CVE Hardened Images & Helm Charts
A from-scratch, security-first replacement for the Bitnami catalog: container images and Helm charts built entirely from source on Wolfi, hardened under a strict 0-CVE build gate, cryptographically signed, and pinned by digest. Free, independent, and fully self-hostable.
Technologies: apko, melange, Wolfi, Helm, Kubernetes, Trivy, Cosign, SLSA, Go, AstroJS
- [Live Site](https://quench-works.com/)
- [ArtifactHub (Verified Publisher)](https://artifacthub.io/orgs/quenchworks)
- [GitHub Organization](https://github.com/quenchworks)
- [Charts](https://github.com/quenchworks/charts)
- [Common (library chart)](https://github.com/quenchworks/common)
- [Website](https://github.com/quenchworks/website)
### sysdesign, System Design Knowledge for AI Agents
A Claude Code plugin that wires tradeoff-first system design knowledge into your AI agent: one skill, eleven commands, and fourteen self-contained reference files. Explain a concept, compare options, pressure-test an architecture, estimate capacity, or prep an interview, with every tradeoff stated, not hand-waved. Original prose, MIT-licensed, works fully offline.
Technologies: Claude Code, Markdown, Python, Mermaid
- [Live Site](https://sysdesign.mkabumattar.com/)
- [View Source](https://github.com/mkabumattar/sysdesign)
### Mathematics - Formula Reference
A clean, searchable web reference for mathematical formulas across algebra, geometry, trigonometry, and calculus. It cuts the clutter most references bury you in. Formulas render with server-side KaTeX on focused topic pages, with no accounts, tracking, or ads.
Technologies: AstroJS, TypeScript, Tailwind CSS, React, KaTeX, MDX
- [Live Documentation](https://mathematics.mkabumattar.com/)
- [View Source](https://github.com/MKAbuMattar/mathematics)
### NetCalc Pro
NetCalc Pro Cloud Engineering Suite is a web application for subnetting, CIDR calculations, and VLSM. It has a Zero Trust, no-backend architecture with a Terminal Brutalist UX, so network engineers can plan address space, share state via Base64 URLs, and use the whole toolset without a backend server.
Technologies: AstroJS, TypeScript, Tailwind CSS, Nanostores
- [Live Documentation](https://netcalc-pro.mkabumattar.com/)
### Rawi - AI CLI Documentation Tool
A CLI tool that generates documentation for command-line applications using AI. Rawi analyzes your CLI commands and creates detailed, structured documentation automatically.
Technologies: TypeScript, Node.js, AI/ML, CLI, NPM
- [Live Documentation](https://rawi.mkabumattar.com/)
- [View Source](https://github.com/withrawi/rawi)
- [NPM Package](https://www.npmjs.com/package/rawi)
### AWS Icons - AWS Architecture Icons for Every Stack
The official AWS Architecture Icons packaged as 16 npm packages under one scope: raw SVG, React, Preact, Vue, Solid, Svelte, Astro, Angular, Lit, Web Components, Alpine.js, htmx, React Native, Qwik, Iconify collections, and SVG sprites. A monthly pipeline syncs the official AWS icon set and releases every package automatically. Supersedes the archived aws-icons and aws-react-icons packages.
Technologies: TypeScript, SVG, React, Vue, Angular, AstroJS, NPM, GitHub Actions
- [Docs & Gallery](https://aws-icons.mkabumattar.com/)
- [View Source](https://github.com/MKAbuMattar/aws-icons)
- [All Packages (@aws-icons)](https://www.npmjs.com/org/aws-icons)
### Devicons Pack - Programming Icons for Every Stack
Devicons (programming languages, tools, and tech logos) packaged as 16 npm packages under one scope, from raw SVG to React, Vue, Svelte, Angular, React Native, Qwik, Iconify, and more. Weekly pipeline tracks devicons/devicon (master feeds stable, develop feeds the beta channel). Supersedes the archived devicons-react package (8k+ weekly downloads).
Technologies: TypeScript, SVG, React, Vue, Svelte, AstroJS, NPM, GitHub Actions
- [Docs & Gallery](https://devicons-pack.mkabumattar.com/)
- [View Source](https://github.com/MKAbuMattar/devicons-pack)
- [All Packages (@devicons-pack)](https://www.npmjs.com/org/devicons-pack)
### Fluent UI Emoji - Microsoft Emoji for Every Stack
Microsoft Fluent UI Emoji packaged as 16 npm packages under one scope: raw SVG, React, Preact, Vue, Solid, Svelte, Astro, Angular, Lit, Web Components, Alpine.js, htmx, React Native, Qwik, Iconify collections, and SVG sprites. A weekly pipeline syncs microsoft/fluentui-emoji and releases every package automatically. Supersedes the archived fluentui-emoji and react-fluentui-emoji packages.
Technologies: TypeScript, SVG, React, Vue, Svelte, AstroJS, NPM, GitHub Actions
- [Docs & Gallery](https://fluentui-emoji.mkabumattar.com/)
- [View Source](https://github.com/MKAbuMattar/fluentui-emoji)
- [All Packages (@fluentui-emoji)](https://www.npmjs.com/org/fluentui-emoji)
### igntui - .gitignore Generator TUI
A Terminal User Interface (TUI) and CLI for generating .gitignore files from gitignore.io templates. It has smart search, multi-template selection, live preview, and caching for instant access to 571+ templates.
Technologies: Python, TUI, CLI, PyPI, curses
- [View Source](https://github.com/MKAbuMattar/igntui)
- [PyPI Package](https://pypi.org/project/igntui/)
### Asciiquarium - Python Edition
An aquarium and sea animation in ASCII art for your terminal. This is a Python reimplementation of the classic Perl asciiquarium, with multiple fish species, sharks, whales, ships, and sea monsters animated at 30 FPS.
Technologies: Python, TUI, ASCII Art, PyPI, Animation
- [View Source](https://github.com/MKAbuMattar/asciiquarium-python)
- [PyPI Package](https://pypi.org/project/asciiquarium/)
### Notebook - Browser-Based Markdown Editor
A simple notebook that lives entirely in your browser. Everything you type is saved automatically and tucked into the URL. Your work stays private and follows you wherever you go, with no account needed.
Technologies: Web, AstroJS, Markdown, Local Storage
- [View Source](https://github.com/MKAbuMattar/notebook)
- [Try Notebook](https://notebook.mkabumattar.com/)
---
# Contact Me
- **Phone**: +962 79 0650 332
- **Email**: info@mkabumattar.com
- **Location**: Amman, Jordan
---
# Cutting a SaaS AWS Bill 41% Without Slowing Delivery
A growing SaaS ran on EKS with a full GitOps pipeline, and it was over its AWS budget nearly every month. The reflex from leadership was the usual one: freeze features until the bill comes down. That would have worked, and it would have been the wrong call. Freezing delivery to save money trades a problem you can measure for one you can't. This is how the bill came down by roughly 41% over two quarters while the team kept shipping through a 25% canary rollout on every release. Almost none of the win came from turning things off in a panic.
The percentages here are representative of what this pattern achieves, not a
single audited client figure. The AWS and Kubernetes mechanics (tagging, Cost
Categories, Budgets, Anomaly Detection, Savings Plans, EKS node groups,
Karpenter, Argo CD, Argo Rollouts) are exactly as described. Your real savings
depend on how much waste you start with and how much of your compute is
commitment-eligible.
## Impact
The bill came down without a feature freeze. Over two quarters the monthly AWS spend dropped by roughly 41% while the team kept shipping through the same 25% canary it uses for any feature.
None of it came from a panic switch-off. Every lever was reversible and canary-guarded, so the savings held without trading away reliability or delivery speed, and cost turned into a normal signal that shows up in pull requests instead of a quarterly fire drill.
## The problem
The bill was growing faster than revenue, which is the signal that matters. Nobody could say where the money went, because nothing was labeled. A single line item for "EC2" across a dozen teams and six node groups tells you nothing you can act on. When finance asked engineering to explain a spike, the honest answer was "we're not sure," and that answer is what turns a cost conversation into a feature freeze.
There was also a dashboard, and everyone pointed at it as proof they were "doing FinOps." A dashboard shows you the number. It does not change anyone's behavior, and it definitely does not tell an engineer that the node group they oversized last sprint is the reason the graph bent upward. Visibility without attribution is just a prettier version of not knowing.
The last piece was fear. Every proposed saving came with "will this break production?" and without data nobody could answer, so nothing happened. The goal was to make cost a normal, reversible engineering decision instead of a quarterly emergency, and to do it without touching the delivery pipeline the team depended on.
## Constraints
A few limits shaped the whole approach.
- **No feature freeze.** Delivery velocity was the business. Any optimization that slowed shipping was off the table.
- **No reliability regressions.** Saving money by removing redundancy or headroom was not a real saving.
- **Keep the delivery model.** Full GitOps, dev to staging to production promotion, and 25% canary rollouts all had to stay exactly as they were.
- **Data stays isolated.** The data tier runs in subnets with no internet route, and that boundary was non-negotiable.
- **Respect confidentiality.** Real dollar figures stay private, so success is reported as percentages and unit economics.
## Architecture
The SaaS runs entirely on a single EKS cluster per environment, across three availability zones. Each VPC has three tiers of subnets: three public subnets for ingress and NAT, three private subnets for the EKS worker nodes, and three isolated subnets with no internet route for the data tier. Everything the product needs runs in the cluster.
_EKS platform architecture_
Compute is split into purpose-built managed node groups rather than one big pool, which is what makes both scheduling and cost control tractable. There are separate node groups for the frontend, the backend, data and ETL work, observability, the GitOps controllers, and the internal dashboards, plus a Spot-backed group for batch and preview workloads. Karpenter handles just-in-time node provisioning on top, so capacity follows demand instead of sitting idle.
The whole platform toolchain lives in the cluster, scheduled onto those node groups: Keycloak for single sign-on across the dashboards, Argo CD, and Grafana; SonarQube and Trivy in the delivery path; Argo CD for GitOps and Argo Rollouts for progressive delivery; and Prometheus with Grafana for observability.
The data tier (RDS with a multi-AZ standby, plus ElastiCache) lives in the
isolated subnets. Those subnets have no NAT and no internet gateway route, so
the databases cannot reach the internet and the internet cannot reach them.
Nodes talk to them only over private VPC routes.
Cost work fails when it lives only in a finance spreadsheet, so the first move was to make spend attributable. A small, enforced tag taxonomy flows into AWS Cost Categories, which maps raw line items to teams and products, and from there into Budgets, Cost Anomaly Detection, and per-team showback.
_Cost attribution and guardrail flow_
The taxonomy was deliberately small. Five mandatory tags, not twenty, because a taxonomy nobody follows is worse than none. `Environment`, `CostCenter`, `Application`, and `Owner` covered almost every question we needed to answer, and `ManagedBy` flagged anything created by hand instead of through code. On EKS those tags also propagate to node groups and volumes, so cluster compute is attributable per team, not lumped under one anonymous bill.
Enforcement matters more than intent, so the tags were governed centrally. AWS Organizations Tag Policies defined the allowed keys and values, Service Control Policies blocked non-compliant resources, and consolidated billing plus the Cost and Usage Report gave one clean view across every account.
_Governance and reporting topology_
Chargeback is tempting, but it needs near-perfect tagging and it starts turf
wars early. We started with showback: show each team its own spend, let
central finance keep paying the bill, and move to chargeback only once the
tags were trustworthy. Awareness drove most of the savings before any money
changed hands internally.
## Delivery: GitOps and progressive rollout
None of the cost work was allowed to disturb delivery, so it helps to see what delivery looks like. Every change runs through the same GitOps pipeline. CI builds and tests, SonarQube enforces a quality gate, and Trivy scans the image and the IaC. Only a clean build pushes to ECR and bumps the image digest in the GitOps manifests repository. Argo CD notices the change and syncs it to the cluster.
_GitOps delivery pipeline_
Nothing goes fully live at once. Argo Rollouts takes over at the cluster and shifts traffic in steps, starting at 25%, then pausing to check analysis metrics before it widens. If the metrics stay healthy it promotes; if they degrade it rolls back on its own, with no human in the loop. That single behavior is what let the cost changes ship safely, because a right-sized deployment that misbehaved would be caught at 25% of traffic, not 100%.
Environments follow the same path every time. A feature or preview environment spins up per pull request on the Spot-backed node group, merges auto-deploy to dev, the same image digest promotes to staging for integration tests, and only then does it reach production behind the canary.
_Environments and promotion_
The canary itself is a few lines of Argo Rollouts config, and it is the same for a feature change or a cost change.
```yaml title="rollout.yaml"
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: frontend
spec:
strategy:
canary:
steps:
- setWeight: 25
- pause: {duration: 10m}
- setWeight: 50
- pause: {duration: 10m}
- setWeight: 100
```
## Implementation
The baseline was code. Rather than tag resources by hand, every provider inherited a default set of tags, so new infrastructure was attributable from the moment it existed.
```hcl title="provider.tf"
provider "aws" {
region = "us-east-1"
default_tags {
tags = {
Environment = "Prod"
CostCenter = "1001"
ManagedBy = "Terraform"
}
}
}
```
Then came the guardrails, automated so nobody had to remember to check a dashboard. A monthly budget with a forecast alert catches planned overspend before the month ends, and Cost Anomaly Detection catches the surprise 3am spike.
```bash title="Budget + anomaly guardrails"
aws budgets create-budget \
--account-id 111122223333 \
--budget '{
"BudgetName": "MonthlyCost",
"BudgetLimit": { "Amount": "50000", "Unit": "USD" },
"TimeUnit": "MONTHLY",
"BudgetType": "COST"
}'
aws ce create-anomaly-monitor \
--anomaly-monitor '{
"MonitorName": "CoreServices",
"MonitorType": "DIMENSIONAL",
"MonitorDimension": "SERVICE"
}'
```
With attribution and guardrails in place, the actual optimization ran as normal GitOps changes, each behind the canary, each with a one-commit rollback.
1. **Find the waste.** Use Cost Explorer grouped by the new tags, plus Kubernetes right-sizing signals from the metrics stack, to rank the most over-provisioned node groups and workloads.
2. **Right-size node groups in reversible steps.** Drop one instance size or one replica at a time, ship it through the 25% canary, and watch the SLOs. The old manifest is one revert away.
3. **Let Karpenter consolidate.** Enable consolidation so underused nodes are drained and replaced with fewer, better-packed ones, and move interruptible and preview work to the Spot node group.
4. **Put dev and staging to sleep.** Scale non-production node groups to zero overnight and on weekends. Nothing runs when nobody is working.
5. **Commit last, not first.** Only after cluster usage was stable did we buy Compute Savings Plans, so we committed to real baseline usage rather than to waste.
The commitment step is where teams most often lose money, by chasing the deepest discount for a workload they are about to change. The rule we used was simple: match the commitment to the roadmap, not to the current instance.
_Choosing the right commitment_
The team was midway through moving several backend services to Graviton for better price-performance. A three-year EC2 Instance Savings Plan on the old family would have looked cheaper on paper and then stranded the moment those services migrated. A Compute Savings Plan gave up a few points of discount but stayed flexible across families, regions, Fargate, and Lambda, which matters even more on EKS where node groups change shape often. That small premium was cheap insurance.
## Results
Over two quarters, the monthly bill came down by roughly 41%, and delivery never paused. Every cost change went out through the same 25% canary as any feature. The breakdown, as representative shares of the total reduction, looked like this.
| Lever | Share of the saving | Nature of the change |
| :------------------------------------------------- | :------------------ | :------------------------- |
| Right-sizing node groups + Karpenter consolidation | Largest | Reversible, canary-guarded |
| Scheduling dev and staging to sleep | Large | Fully reversible |
| Compute Savings Plans matched to the roadmap | Meaningful | 1-year, flexible |
| Spot for batch, preview, and CI | Meaningful | Interruption-tolerant only |
| Storage cleanup and lifecycle policies | Smaller | One-time plus ongoing |
The more durable result was cultural. Cost stopped being a quarterly fire drill. Teams could see their own node-group spend, cost showed up in pull requests as a normal signal, and the "will this break?" fear faded because every change had a canary and a rollback. Treat the percentage as illustrative and the mechanics as the real deliverable.
## Lessons
Attribution is the whole game. Nothing else worked until spend had an owner, because you cannot optimize a shared cluster you cannot see per team. The five-tag taxonomy, enforced in code and propagated to node groups, paid for itself before a single node was resized.
Progressive delivery is what makes cost work safe. On a normal deploy model, right-sizing production feels risky enough that teams avoid it. With a 25% canary and automated rollback, a bad resize is a non-event, so the team actually did the work instead of flinching.
Commit to usage, not to hope. The most expensive mistake in cloud cost work is a long, rigid commitment bought early to chase a headline discount. Buy commitments after usage is stable, prefer flexibility while the architecture is still moving, and treat the discount rate as secondary to not stranding the plan.
If I did it again, I would wire a pull-request cost estimate in on day one. Putting the number in front of the engineer at the moment they change a manifest moved behavior more than any dashboard did.
## Frequently Asked Questions
> **Why not just freeze features until the bill comes down?**
A freeze trades a measurable problem for an unmeasurable one. You save some money and lose delivery velocity, customer momentum, and team morale, none of which show up cleanly on the bill. Almost all of the saving here came from waste and mismatched commitments, not from doing less, so the freeze would have hurt the business while barely touching the real cost drivers.
> **How do you right-size EKS node groups without causing incidents?**
Treat it like any other change. Drop one size or one replica at a time, ship it through the same 25% Argo Rollouts canary as a feature, and watch the SLOs during the pause windows. Let Karpenter consolidate underused nodes rather than doing it by hand. Because every step is a GitOps commit, the rollback is a one-line revert, so a bad resize is caught at 25% of traffic and reverted, not discovered in a postmortem.
> **Why start the canary at 25% instead of a smaller slice?**
Twenty-five percent is a deliberate balance. It is a big enough slice that real traffic patterns and enough metric volume show up quickly, so the analysis step can make an honest call, but small enough that a bad release only touches a quarter of users before it rolls back. Smaller first steps are reasonable for very high-risk changes, but 25% gave this team fast, trustworthy signal without much blast radius.
> **How does the isolated data tier stay reachable if it has no internet?**
The isolated subnets have no NAT and no internet gateway route, so the databases cannot reach the internet and vice versa. The application nodes in the private subnets reach RDS and ElastiCache over private VPC routes only. Anything the data tier genuinely needs from an AWS service goes through VPC endpoints, which keep that traffic on the AWS network rather than the public internet.
> **Does the GitOps and canary setup make cost work harder?**
It makes it safer, which in practice makes it happen. Every cost change is a normal pull request that flows dev to staging to production behind the canary, with SonarQube and Trivy gates on the way. There is no separate risky "cost project," just ordinary changes with the same guardrails as everything else, so teams approve them quickly.
> **What is the single highest-impact first step?**
Enforced tagging. Until spend is attributable per team, product, environment, and node group, every other optimization is guesswork. A small mandatory taxonomy, applied through Terraform default tags and AWS Organizations tag policies, turns the bill from one opaque number into a map you can act on.
## References
- [Amazon EKS](https://docs.aws.amazon.com/eks/latest/userguide/what-is-eks.html)
- [Karpenter](https://karpenter.sh/)
- [Argo CD](https://argo-cd.readthedocs.io/)
- [Argo Rollouts (progressive delivery)](https://argo-rollouts.readthedocs.io/)
- [Argo CD image rollouts walkthrough](https://medium.com/@anandctx/argocd-image-rollouts-a9d91943195d)
- [Keycloak](https://www.keycloak.org/documentation)
- [SonarQube](https://docs.sonarsource.com/sonarqube-server/latest/)
- [Trivy](https://trivy.dev/)
- [AWS Cost Categories](https://docs.aws.amazon.com/cost-management/latest/userguide/manage-cost-categories.html)
- [AWS Budgets](https://docs.aws.amazon.com/cost-management/latest/userguide/budgets-managing-costs.html)
- [AWS Cost Anomaly Detection](https://docs.aws.amazon.com/cost-management/latest/userguide/manage-ad.html)
- [AWS Savings Plans](https://docs.aws.amazon.com/savingsplans/latest/userguide/)
- [Terraform AWS provider default_tags](https://registry.terraform.io/providers/hashicorp/aws/latest/docs#default_tags)
---
# Building an Internal Developer Platform on Backstage and GitOps
Product teams were spending more time waiting on the platform team than building features. Spinning up a new service meant opening a ticket and waiting for someone to provision a repo, wire up CI, write Kubernetes manifests, and hook up deployment. Each of those handoffs added days. We built an internal developer platform that turned that whole sequence into a self-service golden path: Backstage for the portal and software templates, Git as the single source of truth, and Argo CD reconciling desired state into the clusters. A developer picks a template, fills a short form, and gets a working repo plus a running service, with policy and RBAC acting as guardrails rather than manual gates.
## Impact
The headline change was that creating and shipping a service stopped being a ticket and became a form. New services that used to take teams the better part of a sprint to stand up now scaffold in minutes, and the first deploy happens on merge without anyone from the platform team touching it. The platform team moved from doing one-off deploys to maintaining the templates that everyone else uses.
Adoption is the metric that actually matters here, because a platform nobody uses is just more software to run. Within the first quarter most new services were created through the golden path rather than by hand, which is the signal that the paved road was genuinely easier than going around it. These numbers are specific to this rollout and were measured on our own usage, so treat them as a shape to expect rather than a guarantee.
## The problem
Every new service started the same way: a ticket. The platform team owned the repo templates, the CI config, the base Kubernetes manifests, and the deploy pipeline, so nothing shipped without them in the loop. That made sense when there were a handful of services, but it stopped scaling. The queue grew, context-switching killed the platform team's own roadmap, and product teams learned to batch requests, which made each one bigger and slower.
The deeper issue was that knowledge lived in people's heads and in copy-pasted YAML. Two teams standing up similar services would end up with subtly different setups, because each one copied whatever the last project happened to do. There was no paved road, just a lot of dirt tracks that mostly worked. When something went wrong in one of those setups, debugging it meant reverse-engineering choices nobody remembered making.
We wanted product teams to move without asking permission for routine work, while the platform team kept ownership of what "correct" looks like. That is the tension an internal developer platform exists to resolve.
## Constraints
The platform had to satisfy a few hard constraints, and every design decision came back to them.
- **Self-service by default.** The common case, creating and deploying a service, had to happen with zero tickets and no human in the platform team's loop.
- **Git as the source of truth.** Every change to what runs in a cluster had to be a commit, so we get review, history, and a trivial rollback for free. No `kubectl apply` from laptops.
- **Guardrails, not gates.** Policy and RBAC had to be enforced automatically. A human manually approving routine deploys would just recreate the ticket queue we were killing.
- **Paved road, not a walled garden.** Teams with genuinely unusual needs had to be able to step off the golden path without the platform blocking them, as long as they still passed policy.
## Architecture
The platform is three moving parts wired together by Git. Backstage is the front door, where developers discover services and kick off golden paths. Git holds both application code and the deployment config that describes desired state. Argo CD watches Git and reconciles that desired state into the Kubernetes clusters. Backstage never talks to the clusters to make changes; it only ever writes to Git, which keeps the whole system auditable.
_Platform control plane_
The Backstage catalog models the world as a small set of entities, and understanding those makes the rest of the platform click. A `Template` describes a golden path: its input parameters as a JSON schema, and the steps it runs to scaffold a service. Each template produces a `Component`, which is an actual service owned by a `Group` (a team). Argo CD then manages an `Application` resource that points at the component's config in Git and syncs it to a cluster.
_Catalog and deployment model (class view)_
The reason Git sits in the middle of everything is that it turns two hard problems, auditability and rollback, into one solved problem: version control. Every deploy is a diff you can read, and undoing a bad change is `git revert`, which Argo CD then reconciles back automatically.
## Implementation
The heart of the platform is the scaffolding flow. When a developer picks a template and submits the form, Backstage's scaffolder renders a skeleton from the template's inputs, creates a repository, opens a pull request, and registers the new component in the catalog. That is the moment the ticket used to be filed; now it is a button.
_Self-service scaffolding to deploy (sequence)_
A software template is a `Template` entity plus a skeleton directory. The parameters block is a JSON schema, so Backstage renders it as a validated form for free. The steps block is what runs when the form is submitted.
```yaml title="template.yaml (Backstage software template)"
apiVersion: scaffolder.backstage.io/v1beta3
kind: Template
metadata:
name: node-service
title: Node.js service (golden path)
description: A production-ready Node.js service with CI, Helm, and GitOps wired up.
spec:
owner: group:platform
type: service
parameters:
- title: Service details
required: [name, owner]
properties:
name:
title: Name
type: string
pattern: '^[a-z][a-z0-9-]{2,30}$'
owner:
title: Owning team
type: string
ui:field: OwnerPicker
steps:
- id: fetch
name: Fetch skeleton
action: fetch:template
input:
url: ./skeleton
values:
name: ${{ parameters.name }}
owner: ${{ parameters.owner }}
- id: publish
name: Create repository
action: publish:github
input:
repoUrl: github.com?owner=acme&repo=${{ parameters.name }}
defaultBranch: main
- id: register
name: Register in catalog
action: catalog:register
input:
repoContentsUrl: ${{ steps.publish.output.repoContentsUrl }}
catalogInfoPath: /catalog-info.yaml
```
The skeleton ships the boring, correct defaults so no team has to reinvent them: a `catalog-info.yaml` so the service shows up in the catalog, a CI workflow that builds and signs the container image, a Helm chart, and the Argo CD `Application` that ties it to a cluster. That last file is what turns a repo into something GitOps actually deploys.
```yaml title="argocd-application.yaml (scaffolded into the config repo)"
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: node-service-dev
namespace: argocd
spec:
project: default
source:
repoURL: https://github.com/acme/config
path: apps/node-service/dev
targetRevision: main
destination:
server: https://kubernetes.default.svc
namespace: node-service
syncPolicy:
automated:
prune: true
selfHeal: true
```
Because the deploy config lives in Git and Argo CD self-heals, the platform naturally behaves like a state machine. A service moves from scaffolded, to built, to deployed in dev, through a policy gate, and on to production, and every transition is a commit. Modeling it that way made it obvious where the guardrails belong.
_Service lifecycle (state machine)_
Guardrails are enforced at two layers. RBAC in Backstage and in the clusters decides who can do what, and admission policy with OPA or Kyverno decides what is allowed to run at all. A policy that every workload must set resource limits, for example, is a Kyverno rule that rejects the deploy at admission rather than a checklist item in a review.
```yaml title="require-resource-limits.yaml (Kyverno policy)"
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: require-resource-limits
spec:
validationFailureAction: Enforce
rules:
- name: check-limits
match:
any:
- resources:
kinds: ['Pod']
validate:
message: 'CPU and memory limits are required.'
pattern:
spec:
containers:
- resources:
limits:
memory: '?*'
cpu: '?*'
```
The repo layout keeps application code and deployment config separate, which is a deliberate GitOps choice: app repos change on every feature, config repos change on every deploy, and keeping them apart makes the deploy history readable.
- config/
- apps/
- node-service/
- dev/
- kustomization.yaml
- deployment.yaml
- prod/
- kustomization.yaml
- deployment.yaml
- argocd/
- node-service-dev.yaml
- node-service-prod.yaml
Onboarding an existing service that predates the platform is a short runbook rather than a rebuild.
1. Add a `catalog-info.yaml` to the repo so Backstage discovers and indexes the service.
2. Move its deployment manifests into the config repo under a per-environment path.
3. Add an Argo CD `Application` pointing at that path, starting with automated sync disabled.
4. Compare the live cluster state against Git until the diff is clean, then enable automated sync and self-heal.
## Results
The change people felt first was speed. Creating a new service went from a multi-day, ticket-driven sequence to a form that produces a working repo and a service running in the dev cluster off a single merge. The first deploy happens with no platform-team involvement, which is the whole point. Read and write access to what runs is governed by RBAC and policy, so faster did not mean looser.
Adoption is the result that tells you the platform actually worked, and it climbed quickly once the golden path was clearly easier than the old dirt tracks. Within the first quarter the large majority of new services came through templates rather than by hand. The platform team's own time shifted from doing one-off deploys to improving templates and policy, which compounds: one improvement to a template lands in every service scaffolded after it. These numbers are specific to this rollout and were measured on our own usage; treat them as a shape to expect, not a guarantee.
The other measurable win was consistency. Because every service starts from the same skeleton, the drift between projects that used to make debugging miserable mostly disappeared. When we needed to roll out a change like a new required label or a security default, we updated the template and the policy, and the fleet converged instead of needing a hand-edit per repo.
## Lessons
The most important lesson is that a platform lives or dies by adoption, and adoption is earned by making the paved road genuinely faster than going around it. We resisted the urge to mandate the platform early. Instead we made the golden path the path of least resistance, and teams chose it. A mandate on a platform people dislike just produces malicious compliance.
Guardrails have to be automatic to matter. The first time we let a "quick manual approval" creep into a deploy path, we had reinvented the ticket queue in miniature. Encoding the rule as admission policy, so the platform enforces it without a human, is what kept self-service actually self-service.
Finally, treat the golden path as a product with a small number of well-maintained templates, not a template for every conceivable variation. A handful of paths that cover the common cases well beats a sprawling catalog nobody trusts. Teams with unusual needs step off the road and still pass policy, and that is fine. The goal was never to control every service, only to make the right thing the easy thing.
## Frequently Asked Questions
> **What exactly is a golden path?**
A golden path is the supported, opinionated way to do a common task, like creating a new service, with the boring correct defaults already wired in. In this platform it is a Backstage software template that scaffolds a repo, CI, a Helm chart, and the GitOps config in one step. It is a paved road you are free to leave, not a wall you cannot cross.
> **Why put Git in the middle instead of deploying straight from Backstage?**
Because Git turns auditability and rollback into a solved problem. Every deploy is a reviewable diff with history, and undoing a bad change is a revert that Argo CD reconciles automatically. If Backstage pushed changes directly to clusters, you would lose that trail and have to build approval and rollback yourself.
> **How do guardrails avoid becoming the ticket queue you replaced?**
They are enforced by machines, not people. RBAC decides who can act, and admission policy with OPA or Kyverno decides what is allowed to run, both automatically at deploy time. There is no human in the routine path clicking approve, which is exactly the bottleneck a manual gate would recreate.
> **What happens to teams with genuinely unusual requirements?**
They step off the golden path. The platform does not block a team from writing their own manifests or CI, as long as the result still passes policy at admission. The golden path is the default that covers the common cases well, not a hard requirement for every service.
> **How do you onboard services that existed before the platform?**
Add a catalog-info.yaml so Backstage indexes the service, move its manifests into the config repo, and add an Argo CD Application with automated sync off at first. Once the live state matches Git with a clean diff, turn on automated sync and self-heal. It is an adoption runbook, not a rebuild.
> **What is the single most useful metric for this kind of platform?**
Adoption, specifically the share of new services created through the golden path rather than by hand. A platform nobody uses is just more software to operate. If teams choose the paved road on their own, it means the road is genuinely faster and safer than the alternative, which is the whole point.
> **Do you need Kubernetes to build an internal developer platform?**
No. Backstage and GitOps patterns apply to plenty of deployment targets. Kubernetes happens to pair well with Argo CD's reconcile loop and with admission policy, which is why this platform uses it, but the core idea of self-service scaffolding into Git as the source of truth is portable.
## References
- [Backstage: Software Templates](https://backstage.io/docs/features/software-templates/)
- [Backstage: Software Catalog](https://backstage.io/docs/features/software-catalog/)
- [Argo CD: Declarative GitOps CD for Kubernetes](https://argo-cd.readthedocs.io/en/stable/)
- [Argo CD: Application resource specification](https://argo-cd.readthedocs.io/en/stable/operator-manual/declarative-setup/)
- [Kyverno: Kubernetes-native policy management](https://kyverno.io/docs/)
- [Open Policy Agent (OPA)](https://www.openpolicyagent.org/docs/latest/)
- [Team Topologies: platform as a product](https://teamtopologies.com/key-concepts)
---
# Migrating a Monolith to Kubernetes Without a Big-Bang Cutover
Almost every failed "let's move off the monolith" project shares one detail: the plan was a big-bang cutover. Rewrite in parallel, pick a weekend, flip the switch, and pray. This is the opposite of that. A large application moved onto EKS one service at a time using the strangler-fig pattern, with a routing facade in front, traffic shifting gradually per route, and a working rollback at every single step. The team kept shipping features throughout, and at no point was the whole application in the air.
## Impact
The application reached Kubernetes with no big-bang moment. Every route moved gradually behind a facade with a working rollback, so no single step was ever high stakes and delivery never froze.
The gains were structural rather than a one-time event. Each extracted service got independent deploys and its own scaling, so teams stopped blocking each other and the hot paths no longer forced the whole application to scale with them.
## The problem
The monolith itself was not the enemy. It ran fine, the team knew it, and it paid the bills. The problem was that it had become the bottleneck for everything else. Deploys were all-or-nothing, so one risky change held up every other team's work. Scaling meant scaling the entire application even when only one part was hot. And onboarding a new engineer meant handing them the whole thing at once.
The tempting fix, a full rewrite with a cutover, is where teams get hurt. You freeze features to build the replacement, the replacement drifts from the original as the original keeps changing, and the cutover becomes a single high-stakes event with no safe rollback. If anything goes wrong at 2am on migration night, the only option is a panicked revert of everything.
The goal was to get the benefits of independent services without ever betting the business on one cutover. That means the old and new systems have to run side by side, in production, for as long as it takes.
## Constraints
- **No big-bang cutover.** At no point could correctness depend on a single switch-flip.
- **No feature freeze.** The monolith kept shipping features throughout the migration.
- **A rollback at every step.** Each increment had to be revertible in minutes, not hours.
- **No shared-database free-for-all.** Extracted services own their data; the goal was decoupling, not a distributed monolith on one schema.
- **Prove parity before deleting anything.** Old code stayed until the new path was verified against it.
## Architecture
Before the migration, the shape was familiar: an Application Load Balancer in front of a monolith running across an Auto Scaling group, all talking to one shared relational database.
_Before: monolith on EC2_
The target keeps the monolith running, containerized, inside an EKS cluster, and puts a routing facade in front of everything. The facade is the heart of the pattern. It looks at each request and decides whether that path has been migrated to a new service or still belongs to the monolith. Extracted services get their own data stores; the monolith keeps its shared database until its remaining parts are small.
_After: strangler facade on EKS_
The name comes from the strangler fig, a plant that grows around a tree and gradually replaces it. The new system grows around the monolith, taking over one responsibility at a time, until the original is either gone or small enough to leave alone. Nothing about it requires a dramatic finish.
## The routing facade
The facade is where the safety comes from. Every request enters through it, and a route table decides the destination. A path that has been migrated goes to the new service; everything else defaults to the monolith. Migration of a single route is itself gradual, too: you shift a small percentage of that route's traffic to the new service, watch it, and widen only when it holds. If the new service misbehaves, the facade falls straight back to the monolith, which is still running and still correct.
_Strangler routing_
In practice the facade can be an ingress with weighted routing, an API gateway, or a service mesh. The mechanism matters less than the property: per-path routing plus per-path traffic weight plus instant fallback.
```yaml title="facade-route.yaml (illustrative weighted routing)"
# /users is being migrated: 10% to the new service, 90% still to the monolith.
http:
- match:
- uri:
prefix: /users
route:
- destination: {host: users-service}
weight: 10
- destination: {host: monolith}
weight: 90
- route: # default: everything else stays on the monolith
- destination: {host: monolith}
weight: 100
```
## Implementation
The migration ran as a loop, not a project plan with an end date. Each pass picked one seam, extracted it, shifted traffic, verified, and cleaned up.
_Extraction sequence_
1. **Containerize the monolith first.** Before extracting anything, get the monolith itself running in EKS behind the facade. Now old and new live in the same place, and the facade is the only thing in front.
2. **Pick a loosely-coupled seam.** Choose a capability with a clear boundary and a data set it mostly owns, for example users or billing. Avoid the tangled core on the first pass; early wins build trust.
3. **Build the service with its own data.** Give the extracted service its own database rather than pointing it at the monolith's schema. Backfill and keep it in sync during the transition, but the target is independent ownership.
4. **Route to it gradually.** Add the path to the facade and shift a small slice of traffic, then widen. Watch latency and error rates during each step, and keep the monolith path warm as a fallback.
5. **Verify parity, then delete.** Once the new service matches the monolith's behavior under real traffic, remove that code from the monolith. Deleting the old path is what makes the win permanent.
6. **Repeat, and know when to stop.** Move to the next seam. Stop when what remains is small and stable enough that extracting it would cost more than it returns.
The most dangerous shortcut is pointing a new service at the monolith's
database so you can "extract later." That gives you two services coupled
through one schema, which is a distributed monolith: all of the network
overhead, none of the independence. Give the service its own data, even if
that means a sync period during the transition.
Data is the genuinely hard part, and it is worth being honest about that. Moving stateless request handling is straightforward; moving the data it owns without downtime is not. The workable approach is to give the new service its own store, backfill it, keep it in sync while both paths run, and cut the monolith's write path over only once the new service is authoritative and verified. Where strict consistency is required during the overlap, treat the monolith as the source of truth until the very last step.
Measure the migration by how much of the monolith is gone, not by how many
services exist. A useful signal is the share of production traffic served by
extracted services and the amount of code deleted from the monolith. Creating
services without deleting code from the original is motion without progress.
## Results
The application moved onto EKS without a single cutover event and without a feature freeze. Because each route shifted gradually with a live fallback, no migration step was a high-stakes moment; the riskiest change only ever touched a small slice of one path at a time. Independent deploys arrived for each extracted service, so teams stopped blocking each other, and the hot paths could scale on their own instead of forcing the whole application to scale with them.
The migration also did not finish in the storybook sense, and that was the right outcome. A stable, low-change remainder of the monolith stayed in place, containerized and behind the facade, because extracting it would have cost more than it returned. Treat "the monolith is gone" as a possible ending, not the goal.
## Lessons
The facade is the whole safety story. Because every request always had a valid destination and an instant fallback, no step was irreversible. That single property is what let the team move quickly instead of cautiously.
Extract the easy seams first. The instinct to start with the messy core is a trap. Early, low-risk extractions build the tooling and the team's confidence, so the hard ones later are routine instead of terrifying.
Data ownership is the real migration. The service boundary is easy; the data boundary is the work. Any plan that hand-waves the database is a plan to build a distributed monolith.
Give yourself permission to stop. The goal was never zero monolith. It was independent, deployable, scalable services for the parts that needed it, and a small stable remainder for the parts that did not.
## Frequently Asked Questions
> **What exactly is the strangler-fig pattern?**
It is an incremental migration approach where a new system grows around an old one and takes over its responsibilities one at a time, until the old system is replaced or reduced to a small remainder. A routing facade sits in front and directs each request to either the new component or the old one, so both run in production together and you never need a single cutover.
> **Why not just rewrite and cut over on a weekend?**
Because a cutover is a single high-stakes event with no safe rollback. You freeze features to build the replacement, it drifts from the original as the original keeps changing, and if anything breaks on migration night your only option is reverting everything at once. Strangler-fig keeps the old system live the whole time, so every step is small and reversible.
> **What makes a good first service to extract?**
Low coupling and a clear data owner. Pick a capability with a clean boundary that mostly owns its own data, like users or billing, so you are not untangling shared state on your first attempt. Early, low-risk wins build the tooling and the confidence you will need for the harder seams later.
> **How do you handle the shared database?**
Give each extracted service its own store rather than pointing it at the monolith's schema. Backfill it and keep it in sync while both paths run, then cut the monolith's write path over only once the new service is authoritative and verified. Sharing one database across services is a distributed monolith and defeats the point of the migration.
> **How do you know when the migration is done?**
When the remaining monolith is small and stable enough that extracting more would cost more than it returns. Track the share of production traffic served by extracted services and the amount of code deleted from the monolith. Done does not have to mean zero monolith; a low-change remainder behind the facade is a perfectly good ending.
## References
- [Martin Fowler: StranglerFigApplication](https://martinfowler.com/bliki/StranglerFigApplication.html)
- [Amazon EKS](https://docs.aws.amazon.com/eks/latest/userguide/what-is-eks.html)
- [AWS Prescriptive Guidance: strangler fig pattern](https://docs.aws.amazon.com/prescriptive-guidance/latest/modernization-decomposing-monoliths/strangler-fig.html)
- [Kubernetes Ingress](https://kubernetes.io/docs/concepts/services-networking/ingress/)
- [Database decomposition patterns](https://microservices.io/patterns/data/database-per-service.html)
---
# Multi-Region Active-Active for a Payments API
A payments API that moves real money had been running comfortably in a single AWS region for years. It was reliable until the day it was not: a regional control-plane incident took the whole service offline for a few hours, and there was no second region to fail over to. For most products that is an outage. For a money-movement API it is stuck settlements, angry partners, and a compliance conversation. The mandate that came out of that incident was simple to say and hard to build: survive the loss of an entire region without losing a committed payment or charging anyone twice.
This is the story of taking that API active-active across two regions. The interesting part is not the traffic routing, which is close to a solved problem. The interesting part is the money: making retries safe, keeping two live databases honest, and proving the failover actually works instead of trusting a diagram.
## Impact
Once the second region went live, a full regional failure stopped being an incident and became a drill. During the quarter after cutover the primary region had two brief degradations, and in both cases traffic shifted to the healthy region inside the target window with no customer-visible errors and, most importantly, no duplicate settlements.
The number the finance and risk teams cared about was the double-charge count, and it stayed at zero. That is not because failures stopped happening. It is because every write path was made idempotent and every retry, whether from a client, a load balancer, or a queue redelivery, converges on the same result.
## The problem
A single-region payments API has two failure modes that a diagram tends to hide. The first is total loss of the region, which is rare but catastrophic and completely outside your control. The second, and the one that actually bites during a failover, is the retry storm: when a region gets shaky, every client, proxy, and queue in the system starts retrying, and if those retries are not idempotent, you turn one payment into several.
The business could tolerate a couple of minutes of elevated latency during a failover. It could not tolerate a lost payment that a customer had already seen succeed, and it absolutely could not tolerate charging a card twice. So the real problem was not "run in two regions." It was "make every money-touching operation safe to repeat, then run in two regions."
## Constraints
The design had to fit inside some hard limits. Payments are regulated, so data residency rules meant certain records could not leave their region of origin, which ruled out a naive single global write master. The team ran on Kubernetes (EKS) and Aurora PostgreSQL already, so the solution had to build on those rather than introduce an exotic new datastore. And the failover had to be measurable: leadership wanted a specific RTO and RPO written down and proven, not a hand-wave.
There was also a people constraint. On-call engineers needed a failover they could trust at 3am without a runbook full of manual database promotion steps, because manual steps under pressure are how a recoverable incident becomes a data-loss incident.
## Architecture
Both regions run the full stack and take live traffic. Route 53 uses latency-based routing with health checks so users hit the closest healthy region, and it fails a region out automatically when its health check trips. Each region has its own ALB, API pods on EKS, an idempotency store, and a database.
_Active-active across two AWS regions_
Two decisions carry the whole design. The idempotency store is a DynamoDB global table, replicated multi-active across both regions, so a key claimed in one region is visible in the other within about a second. The system of record is Aurora Global Database: a writer in the primary region with sub-second physical replication to the secondary, where a reader can be promoted to writer during a failover in roughly a minute. Committed transactions replicate fast enough that the recovery point stays effectively at zero for anything the customer already saw succeed.
The failover path is deliberately boring. A health check trips, DNS shifts, Aurora promotes the secondary, and in-flight retries replay with their idempotency key.
_The regional failover timeline (RTO and RPO)_
## Implementation
The idempotency key does the real work here. Every payment request carries an `Idempotency-Key` header. Before doing any work, the API does a conditional write into the idempotency store to claim that key. If the key already exists, the stored result is returned as-is and no charge happens. If the claim succeeds, the API runs the charge inside a database transaction, records the result under the key with a TTL, and returns it. Any retry, from any region, with the same key gets the same answer.
_An idempotent charge that survives failover_
The subtle bug to avoid is claiming the key and then crashing before the result is stored, which would leave a claimed-but-unfinished key that blocks the retry forever. The fix is to store an in-progress marker at claim time and let the retry either return the finished result or safely resume, with the transaction as the source of truth for whether the money actually moved.
Events flowing out to downstream systems (ledgers, notifications) use a transactional outbox, written in the same transaction as the payment, so an event is emitted exactly once per committed payment and consumers dedupe on the same key. That keeps the two regions from emitting conflicting events for the same operation.
## Results
Across the first quarter live, the primary region degraded twice. Both times Route 53 shifted traffic and Aurora promoted the secondary well inside the two-minute RTO target, and customers saw a short latency bump rather than errors. No payment was lost and nothing was charged twice, which was the entire point.
The less glamorous result was operational confidence. Because the failover is automatic and every write is idempotent, on-call stopped treating a regional wobble as an emergency. The monthly game-day, where a region is deliberately failed out in production-like conditions, went from a nerve-wracking event to a routine check with a green result.
## Lessons
The biggest lesson is that active-active is a data problem wearing a networking costume. Getting traffic to two regions is easy; keeping two live copies of money honest is the hard part, and idempotency is what makes it tractable. If you cannot safely repeat every write, no amount of clever routing will save you during a failover.
The second lesson is that an RTO and RPO you have not tested are just wishes. The game-days repeatedly surfaced small issues (a too-aggressive health-check threshold, a client that did not send idempotency keys on one endpoint) that no diagram would have caught. Failover is a feature, and like any feature it has bugs until you exercise it.
## Frequently Asked Questions
> **Why active-active instead of active-passive?**
Active-passive keeps a warm standby that only takes traffic during a failover,
which means the standby path is rarely exercised and tends to rot.
Active-active runs real traffic through both regions all the time, so the
failover path is the same path you use every day. It costs more, but for a
money-movement API the confidence that the second region actually works is
worth it.
> **How do idempotency keys prevent double charges during a failover?**
Every payment request carries a client-generated key. The API claims that key
in a globally replicated store before charging, and stores the result against
it afterward. If a retry arrives, in the same region or a different one after
failover, the key is already present and the original result is returned
without charging again. The key, not the region, is what guarantees
exactly-once.
> **What is the difference between RTO and RPO here?**
RTO (recovery time objective) is how long the service can be unavailable
before it is back, which here is the couple of minutes it takes DNS to shift
and Aurora to promote a writer. RPO (recovery point objective) is how much
committed data you can lose, which here is effectively zero because
idempotency writes are synchronous and database replication lag stays under a
second for committed transactions.
> **Does data residency break the active-active model?**
It constrains it. Records that legally must stay in their region of origin are
not globally writable, so the design keeps the system of record regional (a
promotable writer per region) rather than a single global write master. The
globally replicated piece is the idempotency store, which holds keys and
results, not the regulated ledger data.
## References
- [Amazon Aurora Global Database](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/aurora-global-database.html)
- [Amazon DynamoDB global tables](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/GlobalTables.html)
- [Making retries safe with idempotent APIs (AWS Builders' Library)](https://aws.amazon.com/builders-library/making-retries-safe-with-idempotent-APIs/)
- [Amazon Route 53 health checks and DNS failover](https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/dns-failover.html)
- [Transactional outbox pattern](https://microservices.io/patterns/data/transactional-outbox.html)
---
# QuenchWorks: Building a 0-CVE Container Image and Helm Chart Catalog
When Bitnami moved its long-trusted catalog behind a paid tier, thousands of teams woke up to a supply-chain problem they didn't choose. The free images they had pinned in production would stop getting updates, and the migration clock started that morning. QuenchWorks is my answer to that. It is a catalog of container images and Helm charts, built entirely from source, hardened under a strict zero-CVE build gate, signed, and free. This is how it's put together and why each decision earns its place.
## Impact
QuenchWorks is real and in production use, not a proof of concept. It replaced Bitnami for common workloads with a catalog that is built from source, provable, and free.
Everything below is verifiable: pull any image and check its signature, SBOM, and provenance yourself. The build gate stays green because the base is small enough that there is almost nothing to be vulnerable in, so the numbers hold instead of drifting the week after launch.
## The problem
The Bitnami catalog was popular for good reasons. It was broad, it was versioned, and it was maintained well enough that most teams never thought about it. Its weaknesses only became obvious once access changed: you didn't control the build, you couldn't prove what was inside a given image, and continued free access was never actually guaranteed.
That last point is the one that bites. When an upstream catalog changes its terms, every `image:` line you pinned becomes a liability at once. You either pay, fork, or scramble. And even before that day comes, an opaque image is its own quiet risk. If you can't see how a layer was produced, you can't reason about what a scanner finds inside it, and you can't answer a security review with anything better than "we trust the vendor."
Most hardened-image alternatives fix one slice of this and charge for the rest. I wanted the whole thing: a catalog broad enough to actually replace Bitnami for common workloads, provable rather than "trust us," and free with no pull limits and no lock-in. If it couldn't be all three, it wasn't worth building.
## Constraints
A handful of hard limits shaped every later decision.
- **Zero fixable CVEs, enforced by the build.** Not a nightly report someone reads later. A gate that fails the build so a vulnerable image never ships in the first place.
- **Built from source.** No repackaging of someone else's opaque binary layers. If it's in the image, we produced it.
- **Provable.** Every image needs a bill of materials and build provenance that a consumer can verify without trusting me.
- **Free to run and maintain.** The whole system builds on free CI, so cost can never be the reason it slips behind a paywall later.
- **Multi-arch.** amd64 and arm64, because production is both now, not one or the other.
Those constraints pull against each other. Zero fixable CVEs across 150+ images sounds impossible if you picture a fat base image. Building everything from source sounds slow. The architecture is what makes them coexist.
## Architecture
The catalog is a pipeline, not a pile of Dockerfiles. Each image is declared as an `apko` plus `melange` spec, built from source on Wolfi, scanned against a zero-fixable-CVE gate, signed, and only then published pinned by digest. Charts sit on a shared library chart and reference those images by digest.
_QuenchWorks build pipeline_
The single decision that makes the zero-CVE gate realistic is the base. QuenchWorks builds on Wolfi, a glibc Linux undistro designed for containers. Most images start with no shell, no package manager, and a tiny set of packages. There's simply very little in the image that can be vulnerable, so keeping the gate green is a fight you can actually win instead of an endless race against a bloated base.
The image itself is assembled declaratively. `melange` builds signed APK packages from source, and `apko` composes those packages plus the Wolfi base into an OCI image with no Dockerfile involved. Because the whole thing is declared, the contents are known, reproducible, and easy to record as a bill of materials.
_Image composition with melange and apko_
The zero-CVE gate and the minimal base are the same decision viewed twice. You
don't reach zero fixable CVEs by patching harder. You reach it by shipping so
little that there's almost nothing to patch.
A catalog is never done, though, because CVEs are disclosed against packages long after an image ships. So the pipeline runs in reverse on a schedule. A nightly Trivy rescan checks every published image, and when a fix lands upstream the affected image rebuilds, re-enters the gate, gets re-signed, and republishes under a new digest. The catalog trends toward zero drift without anyone babysitting it.
_Nightly rescan and self-heal loop_
## Implementation
Each image is a pair of specs. `melange` describes how to build the package from source, and `apko` describes how to assemble the final image. Here's the shape of both, trimmed for clarity.
```yaml title="melange.yaml"
package:
name: my-app
version: 1.2.3
environment:
contents:
packages:
- build-base
pipeline:
- uses: fetch
with:
uri: https://example.com/my-app-${{package.version}}.tar.gz
expected-sha256: '...'
- uses: autoconf/configure
- uses: autoconf/make
- uses: autoconf/make-install
```
```yaml title="apko.yaml"
contents:
repositories:
- https://packages.wolfi.dev/os
packages:
- my-app
- ca-certificates-bundle
accounts:
users:
- username: nonroot
uid: 65532
run-as: 65532
archs:
- x86_64
- aarch64
entrypoint:
command: /usr/bin/my-app
```
The gate is one Trivy call, and it's deliberately strict about what counts. It only fails on CVEs that have a fix available, because a vulnerability with no upstream patch isn't something a rebuild can clear. Everything fixable has to be at zero before the image is allowed out.
```bash title="0-CVE gate"
# Fail the build if any FIXABLE HIGH/CRITICAL vulnerability is present
trivy image --ignore-unfixed --severity HIGH,CRITICAL \
--exit-code 1 ghcr.io/quenchworks/my-app:latest
```
Once an image passes, it gets signed and attested before it's pushed for real. Signing is keyless with Cosign, so there's no long-lived private key to leak, and the SBOM and SLSA provenance ride along as attestations.
1. **Build from source.** `melange` produces signed APKs; `apko` assembles the image with the Wolfi base, nonroot user, and read-only root filesystem defaults.
2. **Gate on zero fixable CVEs.** Trivy scans the full image. One fixable HIGH or CRITICAL fails the pipeline, so a vulnerable image never reaches the registry.
3. **Sign and attest.** Cosign signs the image keyless, then attaches an SPDX SBOM and a SLSA build-provenance attestation.
4. **Publish pinned by digest.** The image is pushed, and the Helm charts reference it by `sha256:` digest, never by a movable tag.
5. **Rescan nightly.** A scheduled Trivy run watches for new fixes and triggers the self-heal rebuild loop.
On the delivery side, charts are the second half of the story. Every chart builds on a shared `quench-common` library chart, so common concerns like security context, probes, and labels live in one place instead of being copy-pasted 120 times. Each chart pins its image by digest.
_Chart topology_
Pinning by digest instead of tag is what makes the catalog trustworthy in practice. A tag can be moved; a digest can't. When a consumer pins a QuenchWorks chart, they get exactly the bytes that passed the gate.
```yaml title="values.yaml"
image:
repository: ghcr.io/quenchworks/postgresql
# Pinned by digest, not tag. This is the exact image that passed the gate.
digest: 'sha256:abc123...'
```
The proof only matters if consumers can check it, so verification gets its own documented step. Anyone can verify an image's signature and attestations before it runs, and an admission policy can enforce that in the cluster so unsigned or unverifiable images never schedule.
```bash title="Verify before you run"
cosign verify \
--certificate-identity-regexp '^https://github.com/quenchworks/' \
--certificate-oidc-issuer https://token.actions.githubusercontent.com \
ghcr.io/quenchworks/postgresql@sha256:abc123...
```
_Consumer verification workflow_
## Results
Everything below is running today, not staged for a launch.
- 150+ container images built from source on Wolfi, each gated to zero fixable CVEs.
- 120+ production Helm charts on the shared `quench-common` library, every one pinned to its image by digest.
- Every image cosign-signed with an SPDX SBOM and SLSA provenance, published under an ArtifactHub verified-publisher organization.
- Multi-arch (amd64 and arm64), nonroot, and read-only root filesystem by default.
- A nightly rescan and self-heal rebuild loop that keeps the catalog current as upstream ships fixes.
- Free and independent. There is no subscription, no registry pull limit, and nothing that locks you in.
The counts above are current catalog figures. Per-image build times and scan times vary by package, so I'd treat any single number there as indicative rather than a benchmark.
## Lessons
The biggest lesson is that the base image choice decides everything downstream. Trying to reach zero CVEs on a fat base is a treadmill; starting from Wolfi's minimal surface turns the gate into something you can keep green for months. If I'd started anywhere else, the self-heal loop would be firing constantly and the whole thing would feel like bailing water.
The second lesson is that provenance costs far less than it's worth. Signing and generating SBOMs added very little build time, but they change the catalog's whole posture. It stops being "trust me" and becomes "verify it yourself," which is the entire point of replacing an opaque upstream. If I were doing it again, I'd wire verification into the consumer docs even earlier, because an unverified signed image is only half the value.
The one thing I'd watch more carefully next time is chart sprawl. The `quench-common` library chart paid for itself immediately, but library conventions need to be locked down early. Once a few charts drift from the shared patterns, every future change gets more expensive.
## Frequently Asked Questions
> **How is 'zero CVE' actually possible across 150+ images?**
It's zero _fixable_ CVEs, and it's mostly a consequence of the base. Wolfi images ship with almost nothing beyond what the app needs, so there's very little surface for a vulnerability to live in. The Trivy gate then fails any build with a fixable HIGH or CRITICAL, so a vulnerable image can't ship. Vulnerabilities with no upstream fix are tracked but don't block, because a rebuild can't clear them.
> **Why pin charts to images by digest instead of a tag?**
A tag is a movable pointer; a digest is the content itself. If you pin `:latest` or even `:1.2.3`, the bytes behind that tag can change. Pinning `sha256:...` guarantees you get exactly the image that passed the gate and was signed. It's the difference between "probably the right image" and "provably the right image."
> **How do I verify an image before running it?**
Use `cosign verify` with the QuenchWorks certificate identity and OIDC issuer, as shown above. That checks the keyless signature against the transparency log. You can also verify the SBOM and SLSA provenance attestations, and enforce all of it in-cluster with an admission policy so nothing unsigned ever schedules.
> **What happens when a new CVE is disclosed after an image ships?**
The nightly Trivy rescan catches it. If the CVE is fixable, the affected image rebuilds from source, goes back through the gate, gets re-signed, and republishes under a new digest. You pick up the fix by moving your pin to the new digest. Nobody has to notice the CVE manually for the loop to run.
> **Is QuenchWorks really free, and what's the catch?**
It's free, with no subscription and no registry pull limits. The catch, if you call it one, is that you verify and pin things yourself rather than outsourcing trust to a vendor relationship. That's a feature for most teams: you get provenance you can audit instead of a support contract you have to believe.
> **Can I use the charts without adopting the whole catalog?**
Yes. The charts and images are independent. You can pull a single hardened image by digest, or install one chart, without buying into everything. The `quench-common` library chart is an implementation detail of the charts, not something you have to adopt in your own repos.
## References
- [QuenchWorks catalog and website](https://quench-works.com/)
- [QuenchWorks on GitHub](https://github.com/quenchworks)
- [Wolfi undistro](https://github.com/wolfi-dev)
- [apko](https://github.com/chainguard-dev/apko) and [melange](https://github.com/chainguard-dev/melange)
- [Trivy vulnerability scanner](https://trivy.dev/)
- [Sigstore Cosign](https://docs.sigstore.dev/)
- [SPDX](https://spdx.dev/) and [SLSA provenance](https://slsa.dev/)
- [ArtifactHub](https://artifacthub.io/)
---
# Zero-Downtime PostgreSQL Major-Version Upgrade at Scale
A multi-terabyte PostgreSQL 12 database was reaching end of life, and the business ran around the clock, so the usual answer of "schedule a maintenance window" was off the table. We upgraded it to PostgreSQL 16 while users kept reading and writing the whole time. The write pause during the final switch was measured in seconds, and the old database stayed hot the entire cutover so we could fail back instantly if anything looked wrong. This is how the migration was designed, rehearsed, and executed.
## Impact
The upgrade happened while users kept reading and writing. The final switch cost a write pause measured in seconds and lost no rows, with the old database kept hot for instant failback.
The safety came from never modifying the source until the very last step, so failback stayed a genuine option the entire time. These numbers are specific to this workload and were measured on our own traffic, so treat them as a shape to expect rather than a guarantee.
## The problem
The classic upgrade paths all needed a window we did not have. `pg_upgrade` with hard links is fast, but it still stops the database, and on a multi-terabyte instance you cannot risk a long tail if something goes wrong mid-upgrade. A dump and restore was measured in hours, which was a non-starter. Even RDS in-place major upgrades take the instance offline for the duration and give you no clean way to abort once they begin.
The workload was write-heavy and latency-sensitive, so we could not just pause the application either. What we needed was a way to build the new version alongside the old one, keep it continuously in sync with live traffic, and switch over in a single short, reversible step.
## Constraints
The migration had to satisfy four hard constraints, and every design decision came back to them.
- **No maintenance window.** The only acceptable interruption was a brief write pause during the final switch, on the order of seconds.
- **Reversible at every step.** Until we were certain, the old PostgreSQL 12 primary had to stay untouched and ready to take traffic back.
- **Multi-terabyte, so the initial copy is not free.** Seeding the target could not lock the source or saturate its I/O during business hours.
- **Correctness of the awkward bits.** Sequences, extensions, large objects, and tables without a primary key all needed explicit handling, because logical replication does not carry all of them for you.
## Architecture
The core idea is a source and a target running side by side, with logical replication streaming changes from the old database to the new one. The application never talks to Postgres directly; it goes through PgBouncer, which is what lets us flip traffic in one place at cutover time.
_Replication and cutover topology_
Logical replication works at the level of rows, not disk blocks, which is exactly why it can span major versions. A publication on the source declares which tables to stream, and a subscription on the target consumes that stream through a replication slot that tracks how far the target has consumed. Those three objects are the whole contract.
_Logical replication objects (class view)_
Physically, nothing about the deployment is exotic. The application nodes pool connections through PgBouncer, PgBouncer points at the source on port 5432, and a second, dormant route to the target waits for the switch. Both databases publish replication lag and error metrics so we can watch the gap close in real time.
_Deployment topology_
## Implementation
The upgrade moved through a fixed set of phases, and treating it as a small state machine kept everyone honest about which step we were on and what "done" meant for each one.
_Upgrade phases (state machine)_
First we turned on logical replication on the source. On RDS that means setting `rds.logical_replication` to `1` in the parameter group and rebooting once, well ahead of the migration. Tables that get updates or deletes need a replica identity so those changes can be matched on the target; a primary key covers most, and anything without one gets `REPLICA IDENTITY FULL`.
```sql title="On the source (PostgreSQL 12)"
-- Stream every table in the app schema.
CREATE PUBLICATION app_pub FOR ALL TABLES;
-- Tables without a primary key need a full replica identity
-- so UPDATE and DELETE can be replicated.
ALTER TABLE audit_events REPLICA IDENTITY FULL;
```
To seed the target without hammering the source during the day, we restored the schema and a recent snapshot into the new PostgreSQL 16 instance first, then created the subscription with `copy_data = false` so it only carried changes from that point forward. On a smaller database you can let the subscription do the initial copy itself, but at multiple terabytes the snapshot route keeps the source calm.
```sql title="On the target (PostgreSQL 16)"
-- Schema and a consistent snapshot are already restored here.
-- Subscribe for ongoing changes only; the bulk data is already present.
CREATE SUBSCRIPTION app_sub
CONNECTION 'host=source.internal dbname=app user=repl'
PUBLICATION app_pub
WITH (copy_data = false, create_slot = true, slot_name = 'app_sub_slot');
```
From there it was a waiting game while the target caught up. We watched the lag from both ends until the gap held near zero under normal write load.
```sql title="Watching the gap close"
-- On the source: how far behind is the subscriber?
SELECT slot_name, active,
pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS behind
FROM pg_replication_slots;
-- On the target: is the subscription streaming and healthy?
SELECT subname, received_lsn, latest_end_lsn, last_msg_receipt_time
FROM pg_stat_subscription;
```
Before touching traffic, we validated parity with dual reads. A read-only checker ran the same queries against both databases and compared row counts and checksums on the busiest tables. We wanted proof that the target was a faithful copy, not just a hopeful one.
Logical replication does not carry everything. Sequences are not replicated, DDL is not replicated, and large objects in `pg_largeobject` are not replicated. Freeze schema changes for the migration window, plan to reset sequences at cutover, and handle large objects separately. `pglogical` can sync sequences for you if you would rather not script it.
The cutover itself was a short, rehearsed sequence. The point of writing it as an ordered exchange between the operator, PgBouncer, and the two databases was to make the timing and the ordering unambiguous.
_Cutover sequence_
We drove it from a runbook with a hard time budget and an explicit rollback branch. If replication did not reach zero lag inside the budget, or the post-cutover smoke tests failed, we resumed writes on the untouched source and walked away to try another day.
_Cutover runbook (activity)_
The two commands at the heart of the switch are pausing the pool and resetting sequences, since sequence values do not come across on their own.
1. Pause new writes at PgBouncer with `PAUSE`, which lets in-flight transactions finish and holds new ones.
2. Confirm on the source that the replication slot has drained to zero lag.
3. Reset every sequence on the target from the source's current values, then run `ANALYZE` so the planner has fresh statistics.
4. Repoint PgBouncer at the target and `RESUME`, so held connections wake up talking to PostgreSQL 16.
5. Run smoke tests. If they pass, announce done. If not, repoint back to the source and resume there.
```bash title="Reset sequences on the target from the source"
# Emit setval() calls from the source, apply them on the target.
psql "$SOURCE" -Atc "SELECT format('SELECT setval(%L, %s);', seqrelid::regclass, last_value)
FROM pg_sequences_lastvals()" \
| psql "$TARGET"
```
## Results
The measured write pause during the switch was a handful of seconds, dominated by draining in-flight transactions rather than any copy. Read traffic was never interrupted, since reads could keep hitting the source until the pool repointed. No rows were lost, which the dual-read parity checks confirmed both before and after cutover.
Because the source stayed primary until the very last step and was never modified, rollback stayed a genuine option right up to the point we chose to decommission it, which we did only after a full business day of clean operation on PostgreSQL 16. These numbers are specific to this workload and were measured on our own traffic; treat them as a shape to expect, not a guarantee.
## Lessons
The single most valuable thing we did was rehearse the cutover against a copy until the runbook was boring. The first rehearsal surfaced the sequence problem, the second surfaced a table with no primary key, and by the third the whole thing was muscle memory.
Watching replication lag as a first-class metric mattered more than any single command. The go or no-go decision at cutover was a number on a dashboard, not a gut feel. And keeping the old primary untouched turned rollback from a scary, multi-hour restore into a one-line repoint, which is what made the whole plan safe enough to run against production in the first place.
## Frequently Asked Questions
> **Why not just use pg_upgrade or an in-place RDS major upgrade?**
Both stop the database for the duration and give you no clean abort once they start. On a multi-terabyte, 24/7 workload that downtime and that lack of a rollback were unacceptable. Logical replication lets you build the new version alongside the old one and switch over in a short, reversible step instead.
> **What is the difference between logical and physical replication here?**
Physical (streaming) replication copies disk blocks and requires both sides to run the same major version, so it cannot help you upgrade. Logical replication ships row-level changes decoded from the WAL, which is version-independent, so a PostgreSQL 12 primary can feed a PostgreSQL 16 subscriber.
> **Why do sequences need special handling?**
Logical replication streams table data but not sequence values, so the target's sequences would still sit at wherever the initial copy left them. If you skip the reset, the first inserts after cutover can collide with existing primary keys. We reset every sequence from the source's live values as part of the cutover, just before resuming writes.
> **How did you seed a multi-terabyte target without hurting the source?**
We restored a recent snapshot and the schema into the target first, then created the subscription with copy_data set to false so it only carried changes from that point forward. That avoids a giant online copy that would compete with production traffic. On a small database you can let the subscription copy the data itself.
> **What was the actual rollback plan?**
Until the final switch, the source stayed primary and unmodified. If replication did not reach zero lag inside the time budget, or the post-cutover smoke tests failed, we repointed PgBouncer back at the source and resumed writes there. Because the source never diverged, that was instant and lossless.
> **Does this work on Amazon RDS?**
Yes. Set rds.logical_replication to 1 in the parameter group and reboot once ahead of time so the source starts producing logical WAL. From there the publication, subscription, and slot work the same as on self-managed PostgreSQL. The target can be a fresh RDS instance on the new major version.
> **What about large objects and extensions?**
Neither rides along automatically. Extensions must be installed on the target before you subscribe, and large objects stored in pg_largeobject are not replicated by native logical replication, so migrate them separately during the window. If your schema leans heavily on either, factor that into rehearsals.
## References
- [PostgreSQL: Logical Replication](https://www.postgresql.org/docs/current/logical-replication.html)
- [PostgreSQL: CREATE PUBLICATION](https://www.postgresql.org/docs/current/sql-createpublication.html)
- [PostgreSQL: CREATE SUBSCRIPTION](https://www.postgresql.org/docs/current/sql-createsubscription.html)
- [PostgreSQL: Replication Slots and pg_replication_slots](https://www.postgresql.org/docs/current/view-pg-replication-slots.html)
- [pglogical (2ndQuadrant / EDB)](https://github.com/2ndQuadrant/pglogical)
- [Amazon RDS for PostgreSQL: Logical Replication](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/CHAP_PostgreSQL.html#PostgreSQL.Concepts.General.FeatureSupport.LogicalReplication)
- [PgBouncer: PAUSE and RESUME](https://www.pgbouncer.org/usage.html)
---
# Content sections
# Blog Posts
108 entries.
- [Kubernetes Health Probes: Building Self-Healing Applications](https://mkabumattar.com/blog/post/kubernetes-health-probes-self-healing), How Kubernetes liveness, readiness, and startup probes turn your application state into signals the control plane acts on, the misconfigurations that cause cascading outages, and how to wire zero-downtime rollouts with readiness gates and preStop drains.
- [Testing Terraform: Static Analysis, Native Tests, and Terratest](https://mkabumattar.com/blog/post/terraform-testing-terratest-native-tests), A practical testing strategy for Terraform modules: the testing pyramid, tflint static analysis, native terraform test with provider mocking, Terratest integration tests in Go, safe teardown, and a GitHub Actions pipeline with OIDC.
- [tRPC: End-to-End Type-Safe APIs in TypeScript Without Codegen](https://mkabumattar.com/blog/post/trpc-end-to-end-typesafe-apis), How tRPC gives full-stack TypeScript teams end-to-end type safety with no code generation: routers, Zod validation, React Query, auth middleware, the v11 features (FormData, SSE, streaming), and when to pick it over REST or GraphQL.
- [FinOps in Practice: How to Build a Cloud Cost Accountability Culture on AWS](https://mkabumattar.com/blog/post/finops-cloud-cost-accountability-aws), How to run FinOps as a real practice on AWS: a lean tagging taxonomy enforced with Terraform and Organizations, automated budgets and anomaly detection, showback vs chargeback, and matching Savings Plans to your architecture roadmap.
- [QuenchWorks: A Zero-CVE, Built-From-Source Replacement for the Bitnami Catalog](https://mkabumattar.com/blog/post/quenchworks-zero-cve-bitnami-alternative-wolfi), When Broadcom moved the free Bitnami catalog to a legacy tier, thousands of teams lost their supply of maintained, hardened container images overnight. QuenchWorks is my answer: over 150 container images and 120 Helm charts, rebuilt from source on Wolfi, scanned to zero fixable CVEs, cosign-signed, and pinned by digest. Here is why I built it and how it actually works.
- [GitHub Actions Reusable Workflows: Build a Shared CI Library Across All Your Repos](https://mkabumattar.com/blog/post/github-actions-reusable-workflows-shared-ci-library), Simplifying and securing CI/CD at scale with GitHub Actions reusable workflows: the differences between reusable workflows and composite actions, OIDC keyless authentication, versioning strategies, and testing methods like "Patch-on-Test" for a centralized pipeline library.
- [Kubernetes Networking Demystified: CNI Plugins, Network Policies, and Pod-to-Pod Communication](https://mkabumattar.com/blog/post/kubernetes-networking-cni-plugins-policies-guide), A friendly, technical guide to Kubernetes networking. We cover how CNI plugins like Calico and Cilium work, how to write Network Policies, and how to debug those annoying connectivity issues.
- [Service Mesh Deep Dive: Istio vs. Linkerd](https://mkabumattar.com/blog/post/service-mesh-istio-vs-linkerd), Trying to figure out service mesh? This article compares Istio and Linkerd on Kubernetes, looking at how they handle traffic, security, and more. Find out which one might be the best fit for you.
- [GitOps vs. Traditional IaC for Kubernetes: A Comparative Analysis](https://mkabumattar.com/blog/post/gitops-vs-traditional-iac-kubernetes-deployment), Explore GitOps vs. Traditional IaC for Kubernetes. This report compares pull-based GitOps (ArgoCD, Flux) with push-based IaC (Terraform), covering workflows, drift detection, security, and rollbacks. Understand how these approaches manage Kubernetes infrastructure and application configurations for better consistency and reliability.
- [AWS Lambda Observability: Monitoring with CloudWatch, X-Ray, and Datadog](https://mkabumattar.com/blog/post/aws-lambda-observability-guide), How to monitor AWS Lambda functions with CloudWatch, X-Ray, and Datadog: tracking performance, troubleshooting failures, and cutting costs in serverless applications.
- [Designing SLOs and Error Budgets: Your Blueprint for Sustainable Reliability](https://mkabumattar.com/blog/post/designing-slos-error-budgets-reliability-blueprint), How to design SLOs and error budgets that balance shipping speed with reliability: SLIs, monitoring, business goals, and the policies that turn a budget into a real decision tool.
- [Edge Computing: AWS Lambda@Edge vs. Cloudflare Workers. A Practical Guide](https://mkabumattar.com/blog/post/edge-computing-aws-lambda-at-edge-vs-cloudflare-workers-practical-guide), A practical comparison of AWS Lambda@Edge and Cloudflare Workers: cold starts, pricing, language support, and which one fits low-latency web and IoT workloads.
- [Navigating the Future of Cloud with Multi-Cloud IaC: Pulumi and Crossplane](https://mkabumattar.com/blog/post/multi-cloud-iac-pulumi-crossplane-future-cloud), How Pulumi and Crossplane approach multi-cloud Infrastructure as Code: reducing vendor lock-in, and what each tool is actually good at.
- [Building Resilient Systems: Immutable Infrastructure with Packer and Terraform](https://mkabumattar.com/blog/post/immutable-infrastructure-packer-terraform-guide), How immutable infrastructure with Packer and Terraform replaces servers instead of patching them: core concepts, hands-on workflows, and best practices for building resilient, secure systems.
- [Modular Terraform for Scalable Infrastructure as Code](https://mkabumattar.com/blog/post/modular-terraform-scalable-iac-guide), Best practices for modular Terraform: module design, versioning, testing, and state management for consistent, automated cloud infrastructure.
- [Unmasking Hidden Costs: Your Guide to AWS Cost Optimization Cleanup Strategies](https://mkabumattar.com/blog/post/aws-cost-optimization-cleanup-strategies), Practical AWS cost cleanup strategies: find hidden costs, clean up EC2, EBS, S3, and RDS waste, and automate ongoing optimization with Cost Explorer, Cloud Custodian, and Spot Instances.
- [Chaos Engineering: Testing Resiliency with Chaos Monkey and Gremlin](https://mkabumattar.com/blog/post/chaos-engineering-resiliency-testing-monkey-gremlin), How Chaos Engineering builds resilient systems. Chaos Monkey and Gremlin inject faults in AWS and Kubernetes so you find the weak spots before an outage does. Best practices, experiment types, and frequently asked questions for resilient software testing.
- [What's the Deal with Shift-Left Security, and Why Should You Care?](https://mkabumattar.com/blog/post/shift-left-security-sast-dast-sca-cicd), How to implement shift-left security with SAST, DAST, and SCA in your CI/CD pipeline to cut costs and catch issues earlier, with tool examples like SonarQube, Trivy, and OWASP ZAP.
- [Container Image Signing with Cosign: A Hands-On Guide to Secure Your Supply Chain](https://mkabumattar.com/blog/post/container-image-signing-cosign-guide), Secure your software supply chain with Cosign. This hands-on guide covers container image signing, keyless and KMS methods, CI/CD automation, Kubernetes deployment verification, and advanced security best practices.
- [GraphRAG Explained: Building Knowledge-Grounded LLM Systems](https://mkabumattar.com/blog/post/graphrag-explained-building-knowledge-grounded-llm-systems), Think of GraphRAG as the "detective" upgrade for AI. Knowledge Graphs help LLMs connect distant dots, stop hallucinations, and reason through complex data in ways standard RAG just can't.
- [The Resilience of Timbernetes: An Analysis of In-Place Pod Vertical Scaling in Kubernetes 1.35](https://mkabumattar.com/blog/post/kubernetes-1-35-in-place-pod-vertical-scaling-guide), Kubernetes 1.35 "Timbernetes" changes resource management with In-Place Pod Vertical Scaling. This report covers the shift from "restart-to-scale" to a dynamic update model, with a focus on benefits for JVM and stateful services, and explains how the kubelet, VPA, and cgroups v2 work together to enable zero-downtime scaling, improve node utilization, and manage memory shrink hazards.
- [Taming the Chaos: Let's Sort Out Those Flaky CI/CD Pipelines](https://mkabumattar.com/blog/post/troubleshooting-flaky-ci-cd-pipelines), Learn practical strategies for troubleshooting and preventing flaky CI/CD pipelines. Identify common causes, use debugging tools like GitHub Actions and act, and implement best practices for a more stable development workflow.
- [The Democratization of Container Security: Docker Hardened Images](https://mkabumattar.com/blog/post/democratization-docker-hardened-images-container-security), Explore how Docker democratized container security by open-sourcing 1,000+ Hardened Images under Apache 2.0. Learn about distroless containers, 95% smaller attack surfaces, SBOM/VEX integration, and how to migrate to secure-by-default images using multi-stage builds.
- [Database DevOps: Making PostgreSQL and MongoDB CI/CD Feel Natural](https://mkabumattar.com/blog/post/database-devops-ci-cd-postgresql-mongodb), Ready to ditch manual database updates? Learn how Database DevOps and CI/CD can transform PostgreSQL and MongoDB management for faster, more reliable releases.
- [Microsoft's Prompt Orchestration Markup Language (POML): Structuring the Future of AI Interaction](https://mkabumattar.com/blog/post/microsoft-poml-orchestrating-ai-prompts-for-llms), Microsoft's Prompt Orchestration Markup Language (POML) brings structure, semantic tags, and a styling system to prompt engineering for LLMs, with a VS Code extension and Node.js/Python SDKs.
- [Compliance as Code: Making Security Easier with Terraform and InSpec](https://mkabumattar.com/blog/post/compliance-as-code-nist-iso-27001-gdpr-terraform-inspec), How to use Compliance as Code with Terraform and InSpec to automatically follow NIST, ISO 27001, and GDPR, including the benefits, challenges, and real-world uses of this way to handle security and regulatory compliance.
- [Low-Code vs. Custom Code: Let's Talk About Speed and Tech Debt](https://mkabumattar.com/blog/post/low-code-vs-custom-code-speed-tech-debt), How to weigh development speed against technical debt when choosing between low-code and custom code, and when each approach makes sense, especially for internal tools like Retool.
- [AIOps: Making DevOps Even Better with Smart AI Tools](https://mkabumattar.com/blog/post/aiops-enhancing-devops-with-ai), How AIOps applies machine learning to IT operations: predictive analytics, automated incident response, and the tools (Splunk, Moogsoft) doing this today.
- [Centralized Logging with Loki, Grafana, and Fluent Bit: Making Sense of Your Systems](https://mkabumattar.com/blog/post/centralized-logging-loki-grafana-fluent-bit), Setting up centralized logging with Loki, Grafana, and Fluent Bit for Kubernetes and microservices: deployment, configuration, and troubleshooting for a reliable observability setup.
- [Karpenter vs. Cluster Autoscaler on AWS: Picking the Right Tool for Your Kubernetes Scaling](https://mkabumattar.com/blog/post/karpenter-vs-cluster-autoscaler-aws-kubernetes-scaling), Compare Karpenter and Cluster Autoscaler on AWS for Kubernetes scaling. Understand their architecture, performance, cost efficiency, and choose the right tool for your needs.
- [HashiCorp Vault vs. AWS Secrets Manager vs. SOPS: Which One Fits Your Setup](https://mkabumattar.com/blog/post/secrets-management-vault-secrets-manager-sops), Compare HashiCorp Vault, AWS Secrets Manager, and SOPS for secrets management. Understand their features, security, automation, and best practices to choose the right tool.
- [How AI and LLMs Are Changing DevOps Incident Response](https://mkabumattar.com/blog/post/ai-powered-devops-incident-response-llms), How AI and Large Language Models are changing DevOps incident response: practical applications, benefits, challenges, and where this is heading.
- [Platform Engineering: Building Internal Developer Platforms (IDPs)](https://mkabumattar.com/blog/post/platform-engineering-building-internal-developer-platforms), Platform engineering and Internal Developer Platforms (IDPs): the benefits, the common pitfalls, best practices for building one, and how Spotify built Backstage.
- [The Real Talk on Microservices vs. Monoliths](https://mkabumattar.com/blog/post/0076-the-dark-side-of-microservices-when-to-avoid-them), Microservices aren't always the answer. The hidden complexities, the challenges Amazon Prime Video ran into, and when sticking with a monolith is the smarter architectural choice.
- [GitHub Actions vs. GitLab CI for Monorepos: Which One Wins?](https://mkabumattar.com/blog/post/github-actions-vs-gitlab-ci-for-monorepos), Comparing GitHub Actions and GitLab CI for monorepo management. Analyze features, CI/CD pipelines, parallel jobs, caching, secrets management, and real-world experiences to choose the best platform.
- [Full-Stack Observability with OpenTelemetry: Getting a Clear View of Your Systems](https://mkabumattar.com/blog/post/full-stack-observability-opentelemetry), A full view of your complex systems using OpenTelemetry: the basics, the benefits, using it with Prometheus and Grafana, and answers to common questions.
- [Navigating Growth: Building a Secure and Scalable AWS Environment with a Multi-Account Architecture and Control Tower](https://mkabumattar.com/blog/post/multi-account-aws-control-tower), Use a multi-account AWS architecture with AWS Control Tower for stronger security, compliance, and simpler management. Includes best practices and FAQs.
- [Policy as Code with Open Policy Agent: A Technical and Governance Perspective](https://mkabumattar.com/blog/post/policy-as-code-opa-guide), Use Policy as Code with Open Policy Agent (OPA) to boost your cloud governance, security, and compliance. This guide covers the basics, benefits, how to integrate with Terraform and Kubernetes, common challenges, and real-world examples.
- [Zero Trust Architecture in DevOps Pipelines: Secure Your CI/CD Workflows](https://mkabumattar.com/blog/post/zero-trust-devops-pipelines-securing-ci-cd), How to integrate Zero Trust Architecture into DevOps CI/CD pipelines using AWS, IAM, and micro-segmentation, so you can secure deployments without slowing down delivery.
- [10+ Secret Git Commands That Will Save Hours Every Week](https://mkabumattar.com/blog/post/10-secret-git-commands-to-save-time), Discover 10+ secret Git commands that will save hours every week! Learn advanced Git techniques for undoing mistakes, managing commits, automating workflows, and optimizing repositories. Perfect for DevOps and GitHub users.
- [Streamlining GitHub Organization Management with Terraform](https://mkabumattar.com/blog/post/streamlining-github-organization-management-with-terraform), How to manage a large GitHub organization efficiently using Terraform, including the key Terraform resources for automating user access, team structures, and repository configurations.
- [Getting Addicted to Coding: Why We Love Programming More Than Sleep](https://mkabumattar.com/blog/post/getting-addicted-to-coding), Why programming is so addicting and how to keep a healthy balance while pursuing it: the thrill of problem-solving, instant feedback, and practical ways to avoid burnout.
- [Deploying Infrastructure with Terraform in CI/CD Pipelines](https://mkabumattar.com/blog/post/deploying-infrastructure-with-terraform-in-ci-cd-pipelines), How to deploy infrastructure using Terraform in a CI/CD pipeline: where Terraform fits into DevOps workflows and how to build a GitHub Actions pipeline for Terraform automation.
- [AI is Not Real: A Software Engineering Perspective](https://mkabumattar.com/blog/post/ai-is-not-real), Modern AI is not intelligent in the human sense. It is large-scale statistical pattern matching and mathematical optimization. Here is what that means for the systems we build, why probabilistic chains fail, and how hybrid architectures make them reliable.
- [When to Use Serverless?](https://mkabumattar.com/blog/post/when-to-use-serverless), When serverless architecture fits your project, when it does not, and what real-world teams learned running it in production.
- [Why You Should Not Use Else Statements in Your Code](https://mkabumattar.com/blog/post/why-you-should-not-use-else-statements), Why avoiding else statements leads to cleaner, more maintainable code: guard clauses, establishing contracts, adding new conditions without nesting, and when an else is still the right call.
- [How to Avoid Over-Engineering Your Code?](https://mkabumattar.com/blog/post/how-to-avoid-over-engineering-your-code), The causes and symptoms of over-engineering in software development, and practical strategies for keeping code simple and aligned with business needs.
- [Software Engineering Principles Every Developer Should Know](https://mkabumattar.com/blog/post/software-engineering-principles-every-developer-should-know), The software engineering principles every developer should know: DRY, KISS, and YAGNI. What each one asks of you, and Python examples of the same code before and after applying them.
- [Phases of the Modernization Process](https://mkabumattar.com/blog/post/phases-of-the-modernization-process), The key phases of the modernization process: aligning with business goals, understanding dependencies, involving customers, and future-proofing your IT infrastructure.
- [React Context API for State Management](https://mkabumattar.com/blog/post/react-context-api-state-management), A practical look at the React Context API for state management, including how to build a simple shared state system with Next.js and TypeScript.
- [Becoming an AWS Pro: A Deep Dive into Amazon Elastic Container Service](https://mkabumattar.com/blog/post/aws-ecs-deep-dive), A look at Amazon Elastic Container Service: how it compares to EC2 and EKS, its common use cases, and the path from on-premises servers to EC2 instances to ECS Fargate.
- [Building a Code Generative AI Model](https://mkabumattar.com/blog/post/building-a-code-generative-ai-model), How to build a Code Generative AI model as a software engineer: how AI writes code, step-by-step instructions to build your own, and answers to common questions about AI-generated code.
- [Understanding Generative AI in Depth](https://mkabumattar.com/blog/post/understanding-generative-ai-in-depth), A guide to Generative AI for experienced software engineers: what it is, how it differs from traditional AI and machine learning, practical use cases, and answers to frequently asked questions.
- [The ORM Dilemma: To Use or Not to Use](https://mkabumattar.com/blog/post/why-not-to-use-orm-in-nodejs), A look at Object-Relational Mapping (ORM) in Node.js, TypeScript, and Express: the real pros and cons, and when and why to consider alternatives to ORM for your database operations.
- [Caching Strategies with Redis in Node.js and TypeScript](https://mkabumattar.com/blog/post/caching-strategies-with-redis-in-node-js-and-typescript), A look at caching in Redis for Node.js and TypeScript applications: the Cache-Aside, Read-Through, Write-Through, and Write-Behind patterns, plus a practical Redis cache key strategy.
- [Scaling Up, Staying Strong: Hands-On AWS CloudFormation Techniques for Building Resilient and Scalable Systems](https://mkabumattar.com/blog/post/scaling-up-staying-strong-hands-on-aws-cloudformation-techniques-for-building-resilient-and-scalable-systems), Building resilient and scalable systems matters for businesses that need to meet growing demand and maintain high availability. This article covers how AWS CloudFormation, an infrastructure as code tool, helps you build resilient and scalable systems with a hands-on approach.
- [Best Practices for Infrastructure Automation in a Cloud Native AWS Environment](https://mkabumattar.com/blog/post/best-practices-for-infrastructure-automation-in-a-cloud-native-aws-environment), Best practices for infrastructure automation in a cloud native AWS environment, covering security and compliance, CI/CD pipelines, monitoring and auto scaling, performance optimization, and cost management.
- [Orchestrating Infrastructure with Terraform](https://mkabumattar.com/blog/post/orchestrating-infrastructure-with-terraform), Terraform lets you create and manage infrastructure across multiple cloud providers by writing declarative configuration files. This post covers how Terraform provisions AWS resources, from execution plans to state management and reusable modules.
- [Deploying Serverless Applications with AWS SAM](https://mkabumattar.com/blog/post/deploying-serverless-applications-with-aws-sam), AWS SAM (Serverless Application Model) extends AWS CloudFormation with a simpler syntax for defining and deploying serverless functions, APIs, and event sources. This post covers SAM templates, local testing with the SAM CLI, and best practices for building serverless applications on AWS.
- [AWS CloudFormation for Infrastructure as Code (IaC)](https://mkabumattar.com/blog/post/aws-cloudformation-for-infrastructure-as-code-iac), AWS CloudFormation lets you declare your infrastructure as JSON or YAML templates, so you can create, update, and manage AWS resources consistently through code instead of manual console clicks. This post covers templates, stacks, change sets, rollbacks, and best practices for using CloudFormation as part of an infrastructure as code workflow.
- [Cloud Native Infrastructure on AWS](https://mkabumattar.com/blog/post/cloud-native-infrastructure-on-aws), An overview of cloud native infrastructure on AWS: what it means, the core services involved (Lambda, DynamoDB, S3), and the architectural considerations for building scalable, resilient applications.
- [Understanding Infrastructure as Code (IaC)](https://mkabumattar.com/blog/post/understanding-infrastructure-as-code-iac), Infrastructure as Code (IaC) manages infrastructure through code instead of manual configuration. This post covers its benefits, the declarative vs. imperative approaches, and popular tools like Terraform, CloudFormation, and Ansible.
- [Infrastructure Automation on AWS: CloudFormation, SAM, and Terraform for IaC](https://mkabumattar.com/blog/post/infrastructure-automation-on-aws-cloudformation-sam-and-terraform-for-iac), Scalable, resilient systems need infrastructure you can define in code. This post walks through how AWS CloudFormation, AWS SAM, and Terraform automate infrastructure in a cloud native AWS setup.
- [REST API vs RESTful API: Architecture and Constraints Explained](https://mkabumattar.com/blog/post/rest-api-vs-restful-api-architecture-and-constraints-explained), What separates REST API from RESTful API, and what statelessness, uniform interface, client-server separation, cacheability, and layered systems actually mean in practice.
- [TypeScript vs. JSDoc: Static Type Checking in JavaScript Compared](https://mkabumattar.com/blog/post/typescript-vs-jsdoc-exploring-the-pros-and-cons-of-static-type-checking-in-javascript), A comparison of TypeScript and JSDoc for catching type errors in JavaScript: how each one works, what each one costs, and when to pick one over the other.
- [RESTful API vs. GraphQL: Which API is the Right Choice for Your Project?](https://mkabumattar.com/blog/post/restful-api-vs-graphql-which-api-is-the-right-choice-for-your-project), How RESTful and GraphQL APIs compare for a real-time data application, and how to decide between the tried-and-true REST approach and the more flexible GraphQL approach for your project.
- [The AWS Well-Architected Framework Explained](https://mkabumattar.com/blog/post/aws-well-architected-framework-explained), An overview of the AWS Well-Architected Framework and its six pillars: Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimization, and Sustainability, with examples of how organizations apply each one.
- [Get Started with Building ReactJS and Docker: A Complete Guide](https://mkabumattar.com/blog/post/get-started-with-building-reactjs-and-docker-a-complete-guide), This article is a complete guide for building and deploying a React JS application with Docker, including prerequisites, environment setup, building, containerizing, deployment, and best practices. It also includes an example of running a React JS application with Docker Compose.
- [Building a Customizable Image Slider in React Using Hooks, SCSS, and TypeScript](https://mkabumattar.com/blog/post/building-a-customizable-image-slider-in-react-using-hooks-scss-and-typescript), This article will guide you through the process of creating a React slider component using Hooks, SCSS, and TypeScript. By the end of this tutorial, you will have a functional and customizable slider that can be easily integrated into your project.
- [How To Run MySQL in a Docker Container: A Step-by-Step Guide with Customization Tips](https://mkabumattar.com/blog/post/how-to-run-mysql-in-a-docker-container-a-step-by-step-guide-with-customization-tips), How to run a MySQL database in a Docker container: starting a container, connecting to it, customizing its configuration, and using volumes for data persistence, plus tips for optimizing your setup. For developers, sysadmins, and database administrators.
- [How To Install Docker On Linux In 4 Easy Steps](https://mkabumattar.com/blog/post/how-to-install-docker-on-linux-in-4-easy-steps), A step-by-step guide to installing Docker on Linux in 4 steps, from updating the package index to running your first Docker container.
- [How to Avoid Common Cloud Services Mistakes](https://mkabumattar.com/blog/post/how-to-avoid-common-cloud-services-mistakes), This article will cover common mistakes made when implementing cloud services and offer solutions to avoid them. The introduction will explain the importance of avoiding these mistakes. The article will then discuss specific mistakes such as not fully understanding the service, not securing resources, lack of monitoring, no disaster recovery plan, vendor lock-in and lack of compliance, outdated software, scalability issues, cost concerns, lack of exit strategy, migration and backup/recovery plan. The conclusion will summarize key points and provide additional resources.
- [How to Deploy a Spring Boot Application to AWS CloudFormation](https://mkabumattar.com/blog/post/how-to-deploy-a-spring-boot-application-to-aws-cloudformation), This article will guide you on how to deploy a Spring Boot application to AWS CloudFormation. It covers steps from creating a CloudFormation template, packaging the application into a deployable artifact, updating the stack with the new version of the application and discussing about continuous deployment for automatic updates. This guide is intended for developers who have an existing Spring Boot application and want to deploy it on AWS.
- [Introduction to Spring Boot Framework](https://mkabumattar.com/blog/post/introduction-to-spring-boot-framework), Many developers use the Spring Boot framework to build web apps and microservices. It's built on top of the Spring Framework and adds a number of conveniences that make it a popular choice. This post covers what Spring Boot is, why it's useful, and how to create a basic Spring Boot application.
- [How To Setup Bastion Host on AWS using CloudFormation Template](https://mkabumattar.com/blog/post/how-to-setup-bastion-host-on-aws-using-cloudformation-template), Learn how to set up a secure Bastion Host on AWS using CloudFormation templates. This tutorial covers the steps to create a VPC, subnets, security groups, and instances, then test Internet connectivity. Detailed instructions and sample code are included.
- [How To Setup Bastion Host on AWS using AWS CLI](https://mkabumattar.com/blog/post/how-to-setup-bastion-host-on-aws-using-aws-cli), In this post, we will learn the best practices of setting up a Bastion Host on AWS using the AWS CLI for secure and remote access to EC2 instances within a Virtual Private Cloud (VPC). We will guide you through creating a VPC, subnets, internet gateway, and configuring the Bastion Host with the appropriate permissions. This post is intended for those who are familiar with AWS and have some basic knowledge of networking and SSH.
- [How to Setup Jenkins on AWS Using CloudFormation](https://mkabumattar.com/blog/post/how-to-setup-jenkins-on-aws-using-cloudformation), We will be using CloudFormation to set up Jenkins on AWS. CloudFormation is a service that helps you model and set up your AWS resources so that you can spend less time managing those resources and more time focusing on your applications that run in AWS.
- [How to CI/CD AWS With Github using Jenkins](https://mkabumattar.com/blog/post/how-to-ci-cd-aws-with-github-using-jenkins), In this post, I will show you how to set up a CI/CD pipeline using Jenkins and GitHub to deploy a simple PHP application to development and production environments on AWS. With this setup, you can deploy your application to AWS with a single click.
- [How to Install Jenkins on AWS EC2 Instance](https://mkabumattar.com/blog/post/install-jenkins-on-aws-ec2-instance), In this post, I will show you how to create an EC2 instance on AWS and install Jenkins on it.
- [Run TypeScript Without Compiling](https://mkabumattar.com/blog/post/run-typescript-without-compiling), We can run TypeScript without compiling it to JavaScript. This is useful for debugging and testing. In this post, I will show you how to do it.
- [React With Redux Toolkit](https://mkabumattar.com/blog/post/react-with-redux-toolkit), In this post, we will learn how to use Redux Toolkit to manage the state of our React application.
- [What is DevOps?](https://mkabumattar.com/blog/post/what-is-devops), DevOps is a set of practices that combines software development (Dev) and information technology operations (Ops). It aims to shorten the systems development life cycle and provide continuous delivery with high software quality.
- [How To Connect A EBS Volume To An Windows EC2 Instance Using Powershell/GUI](https://mkabumattar.com/blog/post/how-to-connect-a-ebs-volume-to-an-windows-ec2-instance-using-powershell-gui), In this post, we will learn how to connect an EBS volume to a Windows EC2 instance using PowerShell or the Disk Management GUI.
- [How To Connect A Two EC2 Instances Database and Files Transfer Using AWS CLI](https://mkabumattar.com/blog/post/how-to-connect-a-two-ec2-instances-database-and-files-transfer-using-aws-cli), In this post, I will show you how to share a database and files between two EC2 instances using AWS CLI. I will use AWS CLI to create a VPC, EC2 instances, EBS, EFS, and security groups. Then I will replicate the database between the instances and share files between them.
- [Understanding Software Versioning](https://mkabumattar.com/blog/post/how-version-number-software-works), Explore the essentials of software versioning, including semantic versioning, rules, formats, and tools to manage dependencies and releases effectively.
- [How To Connect A Two EC2 Instances Data Transfer Using AWS CLI Without AWS EFS](https://mkabumattar.com/blog/post/how-to-connect-a-two-ec2-instances-data-transfer-using-aws-cli-without-aws-efs), In this post, I will show you how to transfer data between two EC2 instances using AWS CLI, without AWS EFS.
- [How To Create a AWS S3 Bucket Using AWS CLI](https://mkabumattar.com/blog/post/how-to-create-a-aws-s3-bucket-using-aws-cli), In this post, I will show you how to create a AWS S3 bucket using AWS CLI.
- [How To Create a DynamoDB Table Using AWS CLI](https://mkabumattar.com/blog/post/how-to-create-a-dynamodb-table-using-aws-cli), In this article, we will learn how to create a DynamoDB table using AWS CLI. We will also learn how to add items to the table and how to query the table.
- [What is a CI/CD?](https://mkabumattar.com/blog/post/what-is-a-ci-cd), Continuous Integration and Continuous Delivery are two of the most important concepts in DevOps. This article covers what CI/CD is and how it fits into a software development process.
- [Setup Nextjs Tailwind CSS Styled Components with TypeScript](https://mkabumattar.com/blog/post/setup-nextjs-tailwind-css-styled-components-with-typescript), In this post, we will setup Nextjs Tailwind CSS Styled Components with TypeScript.
- [How to Connect to AWS RDS MySQL Database to EC2 Instance With PHP By Using PDO](https://mkabumattar.com/blog/post/how-to-connect-to-aws-rds-mysql-database-to-ec2-instance-with-php-by-using-pdo), In this post, we will learn how to connect to AWS RDS MySQL Database to EC2 Instance With PHP By Using PDO.
- [How to Install and Configure Node.js on EC2 Instance Amazon Linux 2](https://mkabumattar.com/blog/post/how-to-install-and-configure-nodejs-on-ec2-instance-amazon-linux-2), Node.js does not exist in the default Amazon Linux 2 repository. So, we need to add the Node.js repository to the system. In this post, we will learn how to install and configure Node.js on EC2 Instance Amazon Linux 2.
- [How to Create a AWS RDS MySQL Database and Connect to it using MySQL Workbench](https://mkabumattar.com/blog/post/how-to-create-a-aws-rds-mysql-database-and-connect-to-it-using-mysql-workbench), RDS is a managed service that makes it easy to set up, operate, and scale a relational database in the cloud. It provides cost-efficient and resizable capacity while automating time-consuming administration tasks such as hardware provisioning, database setup, patching and backups. It frees you to focus on your applications so you can give them the fast performance, high availability, security, and compatibility they need.
- [How to Run an Apache Web Server Using Docker on an AWS EC2 Instance](https://mkabumattar.com/blog/post/how-to-run-an-apache-web-server-using-docker-on-an-aws-ec2-instance), We will learn how to create an AWS EC2 instance using AWS CLI in this tutorial. We will also discover how to set up an AWS EC2 instance so that it functions with the Apache web server. We will also discover how to set up an AWS EC2 instance so that it functions with WordPress.
- [How To Create An AWS EC2 Instance Using AWS CLI](https://mkabumattar.com/blog/post/how-to-create-an-aws-ec2-instance-using-aws-cli), How to build a VPC, subnets, an internet gateway, a NAT gateway, a route table, and a security group with the AWS CLI, then launch an EC2 instance whose user data installs Apache, PHP, MariaDB, and WordPress.
- [How to Install WordPress on Amazon Linux 2](https://mkabumattar.com/blog/post/how-to-install-wordpress-on-amazon-linux-2), How to install WordPress on Amazon Linux 2 and configure it to run with Apache, PHP, and MariaDB.
- [How to Install PHP and MariaDB on Amazon Linux 2](https://mkabumattar.com/blog/post/how-to-install-php-and-mariadb-on-amazon-linux-2), How to install PHP 7.4 and MariaDB on Amazon Linux 2, get PHP working with the Apache web server, secure the database, and run a small PHP page that connects to it.
- [How to Install Apache Web Server on Amazon Linux 2](https://mkabumattar.com/blog/post/how-to-install-apache-web-server-on-amazon-linux-2), How to install the Apache web server on Amazon Linux 2, open HTTP through firewalld, and serve a simple HTML page from an EC2 instance.
- [How to Install and Setup FireWall on Amazon Linux 2](https://mkabumattar.com/blog/post/how-to-install-and-setup-firewall-on-amazon-linux-2), How to install and configure firewalld on Amazon Linux 2, including zones, services, and ports.
- [Git SSH Keys for GitHub, GitLab, and Bitbucket on Windows](https://mkabumattar.com/blog/post/git-ssh-keys-for-github-gitlab-and-bitbucket-on-windows), Git talks to remotes over HTTPS by default, so it asks for your username and password on every git pull or git push. GitHub, GitLab, and Bitbucket all let Git authenticate over SSH with a public key instead. Here is how to generate a key on Windows, add it to each service, and stop typing credentials for every Git command.
- [Customization Windows Terminal With Starship](https://mkabumattar.com/blog/post/customization-windows-terminal-with-starship), How to customize Windows Terminal with Starship. Windows Terminal is a modern, fast terminal application for users of command-line tools and shells like Command Prompt, PowerShell, and WSL.
- [VIM Cheat Sheet](https://mkabumattar.com/blog/post/vim-cheat-sheet), VIM is a highly configurable text editor available on most Linux distributions. Ideal for editing files via the command line, its modal design offers distinct modes for various tasks. This cheat sheet provides a quick reference for VIM's essential commands and features.
- [How To Create A Custom VPC Using AWS CLI](https://mkabumattar.com/blog/post/how-to-create-a-custom-vpc-using-aws-cli), The example below creates a VPC with an IPv4 CIDR block, a public subnet, and a private subnet, all with AWS CLI commands. Once the VPC and subnets are configured, you can run an instance in the public subnet and connect to it. You can also start an instance in the private subnet and reach it from the instance on the public one.
- [Setting up JWT Authentication in TypeScript with Express, MongoDB, Babel, Prettier, ESLint, and Husky - Part 2](https://mkabumattar.com/blog/post/setting-up-jwt-authentication-in-typescript-with-express-mongodb-babel-prettier-eslint-and-husky-part-2), Setting up JWT Authentication in Typescript with Express, MongoDB, Babel, Prettier, ESLint, and Husky: Part 2.
- [Setting up Node.js, Express, Prettier, ESLint, and Husky application with Babel and TypeScript - Part 1](https://mkabumattar.com/blog/post/setting-up-node-js-express-prettier-eslint-and-husky-application-with-babel-and-typescript-part-1), Setting up Node JS, Express, Prettier, ESLint and Husky Application with Babel and Typescript: Part 1.
- [Setting up Node JS, Express, MongoDB, Prettier, ESLint and Husky Application with Babel and authentication as an example](https://mkabumattar.com/blog/post/setting-up-node-js-express-mongodb-prettier-eslint-and-husky-application-with-babel-and-authentication-as-an-example), Setting up Node JS, Express, MongoDB, Prettier, ESLint and Husky Application with Babel and authentication as an example.
- [Dotfiles: A Git-Based Strategy for Configuration Management](https://mkabumattar.com/blog/post/dotfiles), A Git-based strategy for managing your dotfiles with a bare repository, so your configuration files stay synchronized and secure across every machine you use.
- [Git SSH Keys for GitHub, GitLab, and Bitbucket on Linux](https://mkabumattar.com/blog/post/git-ssh-keys-for-github-gitlab-and-bitbucket-on-linux), By default Git connects to remotes over HTTPS, so it asks for your login and password every time you run a command like git pull or git push. The SSH protocol is the alternative. It lets you connect to a server and authenticate to use its services. GitHub, GitLab, and Bitbucket all let Git connect through SSH rather than HTTPS. Public-key encryption then removes the need to type a login and password for each Git command.
# Case Studies
6 entries.
- [Multi-Region Active-Active for a Payments API](https://mkabumattar.com/case-studies/post/multi-region-active-active-payments), How a money-movement API was taken active-active across two AWS regions with idempotency keys, conflict-free replication, and a tested RTO and RPO, so a full regional outage never double-charges a customer or loses a committed payment.
- [Building an Internal Developer Platform on Backstage and GitOps](https://mkabumattar.com/case-studies/post/internal-developer-platform-backstage-gitops), How golden paths in Backstage, self-service software templates, and Argo CD let product teams create, build, and ship services without filing tickets to the platform team.
- [Zero-Downtime PostgreSQL Major-Version Upgrade at Scale](https://mkabumattar.com/case-studies/post/zero-downtime-postgres-upgrade), How we moved a multi-terabyte PostgreSQL 12 database to 16 with no maintenance window, using logical replication, dual-read validation, and a rehearsed, timed cutover with a real rollback trigger.
- [Migrating a Monolith to Kubernetes Without a Big-Bang Cutover](https://mkabumattar.com/case-studies/post/monolith-to-kubernetes-strangler-migration), Using the strangler-fig pattern to move a large monolith onto EKS service by service, with a routing facade, gradual traffic shifting, and a rollback at every step.
- [Cutting a SaaS AWS Bill 41% Without Slowing Delivery](https://mkabumattar.com/case-studies/post/aws-cost-optimization-saas-case-study), A FinOps case study on a SaaS running on EKS with full GitOps and progressive delivery: how tagging, right-sizing node groups, and Savings Plans matched to the roadmap cut the AWS bill without freezing feature work.
- [QuenchWorks: Building a 0-CVE Container Image and Helm Chart Catalog](https://mkabumattar.com/case-studies/post/quenchworks-zero-cve-catalog), How a from-scratch catalog replaced Bitnami with 150+ container images and 120+ Helm charts built from source on Wolfi, gated to zero fixable CVEs, signed, and pinned by digest.
# Quick Reference Cheatsheets
41 entries.
- [Python Virtual Environments Cheatsheet](https://mkabumattar.com/cheatsheets/python-venv), A practical reference for venv, pip, requirements files, pyenv version switching, and pipx for isolated global tools.
- [GitHub Actions](https://mkabumattar.com/cheatsheets/github-actions), The workflow YAML you write over and over. Triggers, jobs and steps, secrets, matrix builds, caching, artifacts, and reusable workflows in one reference.
- [Linux Networking](https://mkabumattar.com/cheatsheets/linux-networking), The commands you reach for to diagnose and manage Linux networks. ip for interfaces and routes, ss for sockets, dig for DNS, traceroute for paths, and tcpdump for packets.
- [kubectl](https://mkabumattar.com/cheatsheets/kubectl), kubectl is the command-line tool for talking to a Kubernetes cluster. Use it to deploy apps, inspect and manage resources, stream logs, and debug running pods.
- [Terraform Cheatsheet](https://mkabumattar.com/cheatsheets/terraform), Practical Terraform cheatsheet covering the core CLI commands, state management, import, workspaces, and key HCL patterns for variables, locals, outputs, dynamic blocks, and more. Useful for daily IaC workflows.
- [AWS CLI](https://mkabumattar.com/cheatsheets/aws-cli), An expert reference for the AWS CLI covering configuration precedence, EC2 lifecycle control, recursive S3 operations, JMESPath querying, output formatting, and secure SSM sessions.
- [Docker Swarm](https://mkabumattar.com/cheatsheets/docker-swarm), Docker Swarm reference guide covering swarm initialization, node management, services, stacks, overlay networking, secrets, configs, rolling updates, and cluster monitoring.
- [Chef](https://mkabumattar.com/cheatsheets/chef), Chef reference guide covering installation, cookbooks, recipes, resources, knife commands, server management, and automation workflows for infrastructure configuration management.
- [Docker Compose](https://mkabumattar.com/cheatsheets/docker-compose), Docker Compose reference guide covering services, volumes, networks, ports, environment variables, commands, configurations, and container orchestration best practices.
- [Dockerfile](https://mkabumattar.com/cheatsheets/dockerfile), Dockerfile reference guide covering FROM, RUN, COPY, EXPOSE, CMD, ENTRYPOINT, environment variables, build optimization, best practices, and container image construction.
- [PostgreSQL](https://mkabumattar.com/cheatsheets/postgresql), PostgreSQL reference guide covering psql commands, database creation, tables, queries, functions, joins, transactions, indexes, and advanced SQL operations.
- [Redis](https://mkabumattar.com/cheatsheets/redis), Redis reference guide covering commands, data types, keys, strings, lists, sets, hashes, sorted sets, transactions, pub/sub, and caching strategies.
- [Ansible](https://mkabumattar.com/cheatsheets/ansible), Ansible cheatsheet covering playbooks, inventories, roles, tasks, variables, handlers, ad-hoc commands, modules, and configuration options.
- [SSH](https://mkabumattar.com/cheatsheets/ssh), SSH cheatsheet covering OpenSSH client usage, authentication methods, port forwarding, key management, X11 forwarding, and configuration options. Includes real-world examples and security best practices.
- [VS Code](https://mkabumattar.com/cheatsheets/vscode), Complete VS Code keyboard shortcuts reference including command palette, navigation, editing, debugging, multicursor operations, and advanced features for macOS and Windows/Linux
- [Go](https://mkabumattar.com/cheatsheets/go), Go is a statically typed, compiled programming language designed for simplicity, efficiency, and concurrent programming. It's ideal for building fast, scalable server applications and system tools.
- [TOML](https://mkabumattar.com/cheatsheets/toml), TOML (Tom's Obvious, Minimal Language) is a configuration file format designed to be minimal, readable, and unambiguous. It's commonly used for application configuration, package manifests, and data serialization.
- [Markdown](https://mkabumattar.com/cheatsheets/markdown), Markdown is a lightweight markup language designed for creating formatted text using a simple, readable syntax. It's widely used for documentation, READMEs, blogs, and content creation across the web.
- [JSON](https://mkabumattar.com/cheatsheets/json), JSON (JavaScript Object Notation) is a lightweight, text-based data format used for data exchange. It supports objects, arrays, strings, numbers, booleans, and null values, and every mainstream programming language can read it.
- [YAML](https://mkabumattar.com/cheatsheets/yaml), YAML (YAML Ain't Markup Language) is a human-friendly data serialization language commonly used for configuration files, data exchange, and infrastructure-as-code. It emphasizes readability and uses indentation to structure data.
- [RegEx](https://mkabumattar.com/cheatsheets/regex), Regular expressions (regex or regexp) are patterns used to match character combinations in strings. They handle pattern matching, validation, and text processing across many programming languages.
- [JavaScript](https://mkabumattar.com/cheatsheets/javascript), JavaScript is a high-level programming language that powers the web. It supports object-oriented, functional, and event-driven programming styles.
- [Vim](https://mkabumattar.com/cheatsheets/vim), Vim is a highly configurable text editor built to make creating and changing any kind of text very efficient. It is included as "vi" with most UNIX systems.
- [Python](https://mkabumattar.com/cheatsheets/python), Python is an interpreted, high-level programming language known for its readability and simplicity. It supports multiple programming paradigms including procedural, object-oriented, and functional programming.
- [Bash](https://mkabumattar.com/cheatsheets/bash), Bash is a Unix shell and command language written by Brian Fox for the GNU Project as a free software replacement for the Bourne shell.
- [Dart](https://mkabumattar.com/cheatsheets/dart), Dart is a statically-typed, strongly null-safe programming language optimized for building fast, multi-platform applications, with async/await and object-oriented features.
- [Docker](https://mkabumattar.com/cheatsheets/docker), Docker is a containerization platform for building, shipping, and running applications in isolated environments. This cheatsheet covers the core Docker CLI commands.
- [Git](https://mkabumattar.com/cheatsheets/git), Git is a distributed version control system for tracking code changes, collaborating with teams, and managing project history.
- [Tmux](https://mkabumattar.com/cheatsheets/tmux), Tmux is a terminal multiplexer for managing multiple terminal sessions, windows, and panes within a single screen. Commands for session, window, and pane management.
- [Helm](https://mkabumattar.com/cheatsheets/helm), Helm is the package manager for Kubernetes that simplifies deploying, managing, and upgrading applications through reusable charts. This cheatsheet covers the core Helm CLI commands and workflows.
- [Kubernetes](https://mkabumattar.com/cheatsheets/kubernetes), Kubernetes is an open-source container orchestration platform for automating deployment, scaling, and management of containerized applications.
- [Screen](https://mkabumattar.com/cheatsheets/screen), GNU Screen is a terminal multiplexer for managing multiple terminal sessions, windows, and panes within a single screen. It covers commands for session and window management, and for splitting.
- [Curl](https://mkabumattar.com/cheatsheets/curl), cURL is a command-line tool for making HTTP requests, transferring data using URLs, and testing APIs. Essential commands for web development and API testing.
- [Cron](https://mkabumattar.com/cheatsheets/cron), Cron is a time-based job scheduler in Unix/Linux for running scripts or commands periodically. Covers crontab syntax, scheduling patterns, and management commands.
- [AWK](https://mkabumattar.com/cheatsheets/awk), AWK is a text processing language for pattern scanning and data extraction. Used to process text files, extract columns, and perform calculations on text data.
- [Netstat](https://mkabumattar.com/cheatsheets/netstat), Complete netstat reference covering network connections, listening ports, routing tables, and network statistics with practical examples
- [Netcat](https://mkabumattar.com/cheatsheets/nc), Complete netcat reference covering TCP/UDP connections, file transfers, server testing, port scanning, banner grabbing, and network troubleshooting with practical examples
- [Grep](https://mkabumattar.com/cheatsheets/grep), Complete grep reference with pattern matching, regular expressions, flags, context options, and practical examples for searching text files
- [Find](https://mkabumattar.com/cheatsheets/find), Complete find reference with file searching, filtering by type/size/time, permissions, advanced operations, and practical examples for locating files
- [Chmod](https://mkabumattar.com/cheatsheets/chmod), Complete chmod reference with numeric and symbolic permissions, special bits, directory permissions, and practical examples for managing file access
- [Sed](https://mkabumattar.com/cheatsheets/sed), Complete sed reference with substitution, addressing, deletion, insertion, transformations, in-place editing, and real-world examples for text manipulation
# Code Snippets
18 entries.
- [Bash Script Locking: Prevent Concurrent Runs with a PID File](https://mkabumattar.com/codesnippets/post/bash-script-locking-pid-file), A Bash snippet that uses a PID file so only one instance of a script runs at a time. Essential for cron jobs and automation scripts that must not overlap. Includes stale lock detection and cleanup on exit.
- [AWS DynamoDB CRUD Operations in Node.js with the AWS SDK v3](https://mkabumattar.com/codesnippets/post/nodejs-dynamodb-crud-aws-sdk-v3), A practical Node.js snippet covering DynamoDB put, get, update, delete, and query operations using the modern AWS SDK v3. Includes the DocumentClient pattern, single-table design basics, and error handling.
- [Python Async HTTP Requests with aiohttp: Fetch Multiple URLs Concurrently](https://mkabumattar.com/codesnippets/post/python-async-http-aiohttp-concurrent-requests), A Python snippet demonstrating concurrent HTTP requests using aiohttp and asyncio. Covers creating sessions, fetching multiple URLs in parallel, handling errors gracefully, and controlling concurrency with semaphores.
- [Bash Retry Function: Automatically Retry Failing Commands with Exponential Backoff](https://mkabumattar.com/codesnippets/post/bash-retry-function-exponential-backoff), A reusable Bash function that retries any failing command with configurable attempts and exponential backoff. Ideal for wrapping flaky network calls, AWS CLI commands, or deployment scripts in CI/CD pipelines.
- [Node.js Environment Variable Validation with Zod at Startup](https://mkabumattar.com/codesnippets/post/nodejs-env-validation-zod-startup), Stop trusting process.env blindly. This Node.js and TypeScript snippet validates every environment variable at startup with a Zod schema, coerces strings into real types, and refuses to boot on bad config so you catch mistakes before a single request is served.
- [AWS EC2 Instance Management with Boto3: Start, Stop, and Query Instances](https://mkabumattar.com/codesnippets/post/aws-ec2-instance-management-boto3-python), Learn how to automate AWS EC2 instance management using Python and Boto3. This guide covers authentication with IAM roles, starting and stopping instances, using waiters, filtering by tags, running bulk operations, and handling API errors. Practical code examples included for DevOps engineers and cloud developers
- [Redis Caching Patterns: Cache-Aside, Write-Through & Cache Invalidation](https://mkabumattar.com/codesnippets/post/redis-caching-patterns-architecture), Master production-ready Redis caching patterns with practical examples. Learn cache-aside (lazy loading), write-through, consistency patterns, TTL strategies, and cache invalidation techniques to reduce database load and improve application performance.
- [PostgreSQL Query Optimization: Indexes, EXPLAIN ANALYZE & Execution Plans](https://mkabumattar.com/codesnippets/post/postgresql-query-optimization-indexes-explain), Master PostgreSQL query optimization with practical examples. Learn EXPLAIN ANALYZE interpretation, effective index strategies, query rewriting, and connection pooling to identify and fix slow queries in production microservices.
- [Multi-Environment Secret Management with HashiCorp Vault](https://mkabumattar.com/codesnippets/post/hashicorp-vault-multi-environment-secrets), Manage secrets securely across dev, staging, and production with HashiCorp Vault. This snippet demonstrates dynamic secret generation, rotation, and cross-environment secret syncing patterns.
- [Top 7 Open Source OCR Models for Document Processing](https://mkabumattar.com/codesnippets/post/top-7-open-source-ocr-models), The best open source OCR models for converting documents, images, and PDFs to text. Compare olmOCR, PaddleOCR, OCRFlux, and more, with performance benchmarks and implementation examples.
- [Why printf Beats echo in Linux Scripts](https://mkabumattar.com/codesnippets/post/printf-beats-echo-linux-scripts), Why printf is more reliable than echo for output in Linux scripts. The portability problems with echo, what printf gives you instead, and when each command is the right choice.
- [Essential Bash Variables for Every Script](https://mkabumattar.com/codesnippets/post/essential-bash-variables), Master the most useful Bash special parameters and environment variables. Learn how to use $0, $?, $@, $UID, $EUID, and XDG variables to write more reliable and portable scripts.
- [Per-App Shell History for Zsh](https://mkabumattar.com/codesnippets/post/zsh-per-app-history), Keep your Zsh history organized across different terminal applications. This snippet automatically creates separate history files for each terminal emulator you use.
- [Per-App Shell History for Bash](https://mkabumattar.com/codesnippets/post/bash-per-app-history), Keep your Bash history organized across different terminal applications. This snippet automatically creates separate history files for each terminal emulator you use.
- [Optimizing your python code with __slots__?](https://mkabumattar.com/codesnippets/post/python-slots-optimization), Discover how Python `__slots__` can reduce memory usage by up to 40% in data-heavy applications. Perfect for MLOps pipelines and big data processing where millions of objects consume precious memory resources.
- [List S3 Buckets](https://mkabumattar.com/codesnippets/post/python-list-s3-buckets), Automate AWS S3 interactions. This Python snippet uses Boto3 to easily list all S3 buckets in your account.
- [AWS Secrets Manager](https://mkabumattar.com/codesnippets/post/nodejs-aws-secrets-manager), Access secrets securely in your Node.js apps. This snippet demonstrates fetching sensitive data from AWS Secrets Manager.
- [Check S3 Bucket Existence](https://mkabumattar.com/codesnippets/post/bash-s3-bucket-exists), Validate AWS S3 bucket presence in your scripts. This Bash snippet checks if a bucket exists before proceeding with operations.
# DevTips
18 entries.
- [Terraform Workspaces vs. Directory-Based Environments: What Actually Scales](https://mkabumattar.com/devtips/post/terraform-workspaces-vs-directory-environments), Workspaces look like the easy way to split dev, staging, and prod, but they quietly stop scaling. Here is when workspaces bite, why most teams move to a folder per environment, and how to switch without breaking live infrastructure.
- [GitHub Actions Secrets and Environment Variables: Handle Config the Right Way](https://mkabumattar.com/devtips/post/github-actions-secrets-environment-variables-guide), Stop leaking credentials in your workflows. This dev tip shows how to scope GitHub Actions secrets, swap long-lived keys for OIDC, mask sensitive output, and pass config between jobs without it ending up in your logs.
- [Docker Multi-Stage Builds: Smaller, Safer Images for Production](https://mkabumattar.com/devtips/post/docker-multi-stage-builds-smaller-production-images), Ship lean, secure containers by splitting your build from your runtime. This dev tip shows how Docker multi-stage builds drop the compilers and dev dependencies, so your production image is smaller and has far less to attack.
- [ArgoCD GitOps: Sync Kubernetes Deployments Automatically from Git](https://mkabumattar.com/devtips/post/argocd-gitops-kubernetes-deployments-git-sync), Stop running kubectl apply by hand. This dev tip shows how ArgoCD watches a Git repo and keeps your Kubernetes cluster matching it, with automatic sync, self-heal, and drift detection.
- [Kubernetes Namespaces: Organize, Isolate, and Secure Multi-Team Clusters](https://mkabumattar.com/devtips/post/kubernetes-namespaces-organize-isolate-multi-team), Sharing one Kubernetes cluster across teams without the chaos. This dev tip walks through layered namespace isolation: ResourceQuotas, LimitRanges, default-deny NetworkPolicies, and namespace-scoped RBAC, with copy-paste manifests and a Terraform example.
- [Helm Charts: Templating & Multi-Environment Kubernetes Deployments](https://mkabumattar.com/devtips/post/helm-charts-kubernetes-multi-environment), How Helm templates Kubernetes manifests for multi-environment deployments: values overrides per environment, conditional logic, chart dependencies, and GitOps rollout with ArgoCD.
- [Structured Logging & Log Aggregation with ELK Stack](https://mkabumattar.com/devtips/post/structured-logging-elk-stack), Centralized logging for microservices with Elasticsearch, Logstash, and Kibana: structured JSON logging, the Logstash pipeline, Kibana dashboards, alerting rules, and index lifecycle policies for production.
- [Container Image Vulnerability Scanning in CI/CD with Trivy](https://mkabumattar.com/devtips/post/container-image-vulnerability-scanning-trivy), How to automate container image vulnerability scanning in CI/CD with Trivy: installation, severity thresholds, GitHub Actions and GitLab CI integration, policy enforcement, and remediation workflows.
- [Policy-as-Code Governance with OPA/Rego](https://mkabumattar.com/devtips/post/policy-as-code-opa-rego), How Open Policy Agent enforces infrastructure standards automatically: writing policies in Rego, wiring them into Terraform and Kubernetes, and blocking non-compliant changes in CI/CD before they merge.
- [Setting Up GitHub Copilot Agent Skills in Your Repository](https://mkabumattar.com/devtips/post/github-copilot-agent-skills-setup), How to build custom Agent Skills for GitHub Copilot on the agentskills.io open standard: folder structure, SKILL.md configuration, the progressive-disclosure loading model, and enabling the feature in VS Code.
- [7 Reasons Learning the Linux Terminal is Worth It (Even for Beginners)](https://mkabumattar.com/devtips/post/7-reasons-learning-linux-terminal-worth-it-beginners), Seven concrete reasons the Linux terminal is worth learning: commands are easier to remember than they look, text beats clicking through GUI menus, commands stay stable across versions, and scripts can automate what a GUI cannot.
- [Docker Is Eating Your Disk Space (And How PruneMate Fixes It)](https://mkabumattar.com/devtips/post/docker-disk-space-prunemate), Your Docker host is slowly filling up with unused images, orphaned volumes, and stale build cache. Manual cleanup feels risky, and you might accidentally delete the wrong thing. Here's how PruneMate automates Docker maintenance across your home lab with scheduled cleanup, remote host support, and a clean interface that shows exactly what you're deleting before you commit.
- [Understanding Kubernetes Services: ClusterIP vs NodePort vs LoadBalancer](https://mkabumattar.com/devtips/post/kubernetes-services-clusterip-nodeport-loadbalancer), When to use ClusterIP, NodePort, or LoadBalancer for a Kubernetes Service: how each type works, its best-fit use case, and the security and scaling trade-offs of picking the wrong one.
- [Managing Terraform at Scale with Terragrunt](https://mkabumattar.com/devtips/post/terraform-terragrunt-wrappers), How Terragrunt wraps Terraform to remove duplicated backend and provider config across dev, staging, and production: defining shared settings once, overriding per environment, and letting Terragrunt handle state and module dependencies.
- [HashiCorp Pulls the Plug on CDKTF](https://mkabumattar.com/devtips/post/cdktf-deprecation-hashicorp-terraform), HashiCorp just deprecated CDKTF as of December 10, 2025. If you built your infrastructure in TypeScript, Python, or Go to avoid HCL, your options are HCL with OpenTofu or a move to Pulumi, and vendor lock-in just bit again.
- [Tracing Microservices with OpenTelemetry](https://mkabumattar.com/devtips/post/tracing-microservices-opentelemetry), How OpenTelemetry traces a request across distributed services: instrumenting your code, running a collector, and visualizing the resulting spans in Jaeger or Zipkin to find bottlenecks and errors.
- [Organizing Terraform with Modules](https://mkabumattar.com/devtips/post/organizing-terraform-modules), How to split Terraform code into reusable modules for networking, databases, and other common components, stored in a shared repository so multiple projects and environments can pull from the same source instead of copy-pasted configuration.
- [Securing CI/CD with IAM Roles](https://mkabumattar.com/devtips/post/securing-cicd-with-iam-roles), How to scope IAM roles per environment (dev, staging, production) in a CI/CD pipeline so each stage only gets the permissions it needs, cutting the blast radius if credentials leak or a build step misbehaves.
# Flashcards
11 entries.
- [Docker and Container Fundamentals Flashcards](https://mkabumattar.com/flashcards/post/docker-container-fundamentals-flashcards), A spaced-repetition deck covering images, containers, Dockerfiles, volumes, networking, Compose, and registries.
- [Terraform Associate Flashcards (TA-003)](https://mkabumattar.com/flashcards/post/terraform-associate-ta003-flashcards), Full exam coverage for the HashiCorp Certified Terraform Associate (TA-003) exam using spaced repetition. Covers IaC concepts, CLI, HCL, state, modules, backends, and Terraform Cloud.
- [Kubernetes Administrator Flashcards (CKA)](https://mkabumattar.com/flashcards/post/kubernetes-administrator-cka-flashcards), Full exam coverage for the Certified Kubernetes Administrator (CKA) exam using spaced repetition. Covers cluster architecture, workloads, scheduling, networking, storage, security, and troubleshooting.
- [AWS SysOps Administrator Associate Flashcards (SOA-C02)](https://mkabumattar.com/flashcards/post/aws-sysops-administrator-associate-flashcards), Full exam coverage for the AWS Certified SysOps Administrator Associate (SOA-C02) exam using spaced repetition. Covers monitoring, reliability, deployment, security, networking, and cost optimization.
- [LPIC-2 Linux Engineer Flashcards](https://mkabumattar.com/flashcards/post/lpic-2-linux-engineer-flashcards), Full exam coverage for the LPIC-2 Linux Engineer certification (Exam 201 & 202) using spaced repetition. Covers kernel, boot, storage, networking, security, DNS, web, email, and more.
- [AWS Certified Developer - Associate Flashcards (DVA-C02)](https://mkabumattar.com/flashcards/post/aws-certified-developer-associate-flashcards), Full exam coverage for the AWS Certified Developer - Associate (DVA-C02) exam using spaced repetition. Covers Development, Security, Deployment, Troubleshooting, and AWS SDK/CLI/APIs.
- [Red Hat System Administration I Flashcards (RH124)](https://mkabumattar.com/flashcards/post/redhat-system-administration-rh124-flashcards), Full course coverage for Red Hat System Administration I (RH124-9.0) using spaced repetition. Covers CLI, files, users, processes, networking, storage, DNF, systemd, logging, and shell scripting.
- [AWS Solutions Architect Associate Flashcards (SAA-C03)](https://mkabumattar.com/flashcards/post/aws-solutions-architect-associate-flashcards), Full exam coverage for the AWS Certified Solutions Architect Associate (SAA-C03) exam. Covers all four domains Resilient, High-Performing, Secure, and Cost-Optimized Architectures.
- [AWS Cloud Practitioner Flashcards (CLF-C02)](https://mkabumattar.com/flashcards/post/aws-cloud-practitioner-flashcards), Full exam coverage for the AWS Certified Cloud Practitioner (CLF-C02) exam using spaced repetition. Covers Cloud Concepts, Security, Technology, and Billing.
- [AWS Beginner Flashcards](https://mkabumattar.com/flashcards/post/aws-beginner-flashcards), Core AWS concepts across Compute, Storage, Networking, Databases, IAM, Serverless, Monitoring, and High Availability for beginners using spaced repetition.
- [JavaScript Intermediate Flashcards](https://mkabumattar.com/flashcards/post/javascript-intermediate-flashcards), Intermediate and advanced JavaScript concepts for deeper mastery using spaced repetition.
# Glossary Index
7 entries.
- [CI/CD & Automation](https://mkabumattar.com/glossary/post/ci-cd-and-automation), Continuous integration, delivery, and deployment terms: pipelines, runners, artifacts, and release strategies.
- [Kubernetes Advanced](https://mkabumattar.com/glossary/post/kubernetes-advanced), Advanced Kubernetes terms covering scheduling, networking, storage, security, extensibility, and workloads for platform engineers.
- [Cloud Computing on AWS](https://mkabumattar.com/glossary/post/cloud-computing-aws), Essential Amazon Web Services terms covering compute, storage, networking, databases, and security for cloud engineers and architects.
- [Networking Fundamentals](https://mkabumattar.com/glossary/post/networking-fundamentals), Core networking terms every developer and engineer should know, covering IP addressing, DNS, protocols, routing, and the OSI model.
- [Linux Server Administration](https://mkabumattar.com/glossary/post/linux-server-administration), Essential terms every Linux system administrator and DevOps engineer should know.
- [Containers & Kubernetes](https://mkabumattar.com/glossary/post/containers-and-kubernetes), Essential terms every DevOps engineer should know about containers and Kubernetes.
- [DevOps Basics](https://mkabumattar.com/glossary/post/devops-basics), Essential terms every DevOps and cloud engineer should know.
# Quizzes
62 entries.
- [Spring Boot: Java Application Framework Essentials](https://mkabumattar.com/quizzes/post/spring-boot-fundamentals-quiz), Test your knowledge of Spring Boot covering dependency injection, auto-configuration, REST controllers, Spring Data, and production-ready features.
- [Java: Core Language & JVM Fundamentals](https://mkabumattar.com/quizzes/post/java-fundamentals-quiz), Test your knowledge of Java fundamentals covering OOP, the JVM, collections, generics, streams, and core language features.
- [Rust: Ownership, Borrowing & Memory Safety](https://mkabumattar.com/quizzes/post/rust-fundamentals-quiz), Test your knowledge of Rust fundamentals covering ownership, borrowing, lifetimes, traits, pattern matching, error handling, and memory-safe systems programming without a garbage collector.
- [System Design & Architecture: Scalability & Resilience](https://mkabumattar.com/quizzes/post/system-design-architecture-quiz), Master system design: scalability, reliability, performance, trade-offs, distributed systems patterns, and architectural decisions for production systems.
- [Testing Strategies: Unit, Integration, E2E](https://mkabumattar.com/quizzes/post/testing-strategies-quiz), Master testing: unit testing, integration testing, end-to-end testing, mocking, fixtures, coverage goals, and CI/CD integration.
- [TypeScript Advanced: Types, Generics, Utility Types](https://mkabumattar.com/quizzes/post/typescript-advanced-quiz), Master TypeScript: advanced types, generics, utility types, decorators, and type-safe patterns for scalable codebases.
- [Kubernetes Advanced: Production Operations & Scaling](https://mkabumattar.com/quizzes/post/kubernetes-advanced-quiz), Master Kubernetes at scale: stateful applications, networking, storage, security policies, autoscaling, and production troubleshooting.
- [WebAssembly (WASM): Performance & Interoperability](https://mkabumattar.com/quizzes/post/webassembly-wasm-quiz), Master WebAssembly: near-native performance in browser, non-browser use cases, toolchains (Rust, C++), and integration with JavaScript.
- [Multi-Cloud Strategy & Architecture](https://mkabumattar.com/quizzes/post/multicloud-strategy-quiz), Master multi-cloud deployments: AWS, Azure, GCP, vendor lock-in prevention, cost optimization, disaster recovery across clouds.
- [Azure Advanced: App Service, Functions, Container Instances](https://mkabumattar.com/quizzes/post/azure-advanced-quiz), Master Azure at scale: App Service, Azure Functions, Container Instances, Azure SQL, CI/CD pipelines, and enterprise cloud patterns.
- [Advanced Frontend Patterns: State Management & Performance](https://mkabumattar.com/quizzes/post/advanced-frontend-patterns-quiz), Master advanced frontend architecture: state management (Redux, Zustand), performance optimization, code splitting, lazy loading, and responsive design patterns.
- [AWS Advanced: EC2, RDS, Lambda, CloudFormation](https://mkabumattar.com/quizzes/post/aws-advanced-quiz), Master AWS at scale: EC2 instance types, RDS databases, Lambda functions, CloudFormation infrastructure as code, and AWS architecture patterns.
- [Flutter: Cross-Platform Mobile Development](https://mkabumattar.com/quizzes/post/flutter-mobile-quiz), Build native iOS/Android apps with Flutter: widgets, state management, navigation, performance, and publishing to app stores.
- [Production Backend: Scaling, Monitoring & Reliability](https://mkabumattar.com/quizzes/post/production-backend-scaling-quiz), Run production backend systems: scaling patterns, monitoring, error handling, logging, resilience, capacity planning, and incident response.
- [Astro: Building Fast Web Experiences](https://mkabumattar.com/quizzes/post/astro-fundamentals-quiz), Build lightning-fast websites with Astro: island architecture, partial hydration, content collections, integrations, and zero JavaScript by default.
- [Async JavaScript: Promises, Async/Await, Event Loop](https://mkabumattar.com/quizzes/post/async-javascript-promises-quiz), Master asynchronous JavaScript: callbacks, Promises, async/await, event loop, microtasks. Essential for modern backend and frontend development.
- [React Native: Cross-Platform Mobile Development](https://mkabumattar.com/quizzes/post/react-native-mobile-quiz), Build native iOS/Android apps with React Native: components, navigation, state management, performance optimization, and deploying to app stores.
- [Database Design & Patterns](https://mkabumattar.com/quizzes/post/database-design-patterns-quiz), Master database design: normalization, indexing, query optimization, ACID, transactions, sharding, replication, and enterprise patterns for scalable systems.
- [Advanced CSS: Layouts, Modern Features & Performance](https://mkabumattar.com/quizzes/post/advanced-css-layouts-quiz), Master advanced CSS: Grid, custom properties, animations, transforms, performance optimization. Build sophisticated, responsive layouts with modern techniques.
- [Vue.js Fundamentals: Reactive Components & Templates](https://mkabumattar.com/quizzes/post/vuejs-fundamentals-quiz), Master Vue.js basics: reactive data, templates, components, directives (v-if, v-for), event handling, and computed properties. Build interactive UIs with Vue.
- [Responsive Web Design: Mobile-First & Breakpoints](https://mkabumattar.com/quizzes/post/responsive-design-quiz), Master responsive design: mobile-first approach, media queries, flexible layouts, viewport settings, and testing. Build sites that work on all devices.
- [Advanced API Architecture & Design Patterns](https://mkabumattar.com/quizzes/post/advanced-api-architecture-quiz), Master advanced API design: REST, gRPC, webhooks, API versioning, caching strategies, pagination, and enterprise patterns for scalable integrations.
- [Infrastructure Patterns: IaC, Provisioning & Orchestration](https://mkabumattar.com/quizzes/post/infrastructure-patterns-quiz), Master infrastructure patterns: Infrastructure as Code, provisioning, configuration management, orchestration. Design scalable, maintainable infrastructure.
- [Node.js & Express Fundamentals: Building Server Applications](https://mkabumattar.com/quizzes/post/nodejs-express-fundamentals-quiz), Master Node.js and Express basics: server creation, routing, middleware, request/response handling, error handling, and REST API development.
- [React Fundamentals: Components, Hooks & State Management](https://mkabumattar.com/quizzes/post/react-fundamentals-quiz), Master React basics: components, JSX, hooks (useState, useEffect), state management, props, and event handling. Build modern interactive UIs with React.
- [HTML Semantics & Web Accessibility](https://mkabumattar.com/quizzes/post/html-semantics-accessibility-quiz), Master HTML semantics and web accessibility: semantic tags, ARIA attributes, WCAG guidelines, keyboard navigation, screen readers. Build inclusive web experiences.
- [API Security & Authentication: Protecting Your APIs](https://mkabumattar.com/quizzes/post/api-security-authentication-quiz), Master API security fundamentals: authentication, authorization, OAuth2, JWT, API keys, HTTPS, rate limiting, CORS, and common vulnerabilities. Secure your APIs against attacks.
- [Observability Stack: Monitoring, Logging & Tracing](https://mkabumattar.com/quizzes/post/observability-stack-quiz), Master observability fundamentals: monitoring with metrics, structured logging, distributed tracing, and observability practices. Learn to build observable systems and debug production issues.
- [CSS Fundamentals: Selectors, Box Model & Styling](https://mkabumattar.com/quizzes/post/css-fundamentals-quiz), Master CSS fundamentals: selectors, box model, properties, layout basics, and styling techniques. Build the foundation for web design and responsive layouts.
- [GraphQL Fundamentals: Queries, Mutations & Schema Design](https://mkabumattar.com/quizzes/post/graphql-fundamentals-quiz), Master GraphQL foundations: queries, mutations, subscriptions, and schema design.
- [GitOps at Scale: ArgoCD, Flux & Multi-Cluster Deployments](https://mkabumattar.com/quizzes/post/gitops-at-scale-quiz), Master GitOps patterns for multi-cluster Kubernetes deployments using ArgoCD and Flux v2.
- [Istio: Service Mesh Fundamentals](https://mkabumattar.com/quizzes/post/istio-service-mesh-fundamentals-quiz), Master the fundamentals of Service Mesh, traffic routing, and mutual TLS security within a Kubernetes environment.
- [Elasticsearch: Full-Text Search Engine Fundamentals](https://mkabumattar.com/quizzes/post/elasticsearch-search-engine-fundamentals-quiz), Test your knowledge on full-text search, document indexing, and log analysis using Elasticsearch and Kibana.
- [MongoDB: NoSQL Database Fundamentals](https://mkabumattar.com/quizzes/post/mongodb-nosql-fundamentals-quiz), Master NoSQL fundamentals, BSON document structure, and the MongoDB aggregation pipeline.
- [SQL: Query Fundamentals & Database Concepts](https://mkabumattar.com/quizzes/post/sql-query-fundamentals-quiz), Master the fundamentals of Relational Databases, from basic SELECT statements to complex JOIN operations and database normalization.
- [Go: Programming Fundamentals & Concurrency](https://mkabumattar.com/quizzes/post/golang-programming-fundamentals-quiz), Go runs much of the cloud-native toolchain. Test your knowledge on Go syntax, types, and concurrency patterns.
- [Redis: In-Memory Caching & Data Structures](https://mkabumattar.com/quizzes/post/redis-caching-fundamentals-quiz), Master the fundamentals of Redis, including in-memory data structures, caching patterns, and high-performance messaging.
- [Nginx: Web Server Fundamentals](https://mkabumattar.com/quizzes/post/nginx-web-server-fundamentals-quiz), Test your knowledge of Nginx configuration, reverse proxying, load balancing, and SSL/TLS termination.
- [Packer: Infrastructure Image Building Fundamentals](https://mkabumattar.com/quizzes/post/packer-image-building-fundamentals-quiz), Master automated machine image creation with this quiz on Packer builders, provisioners, and HCL templates.
- [HashiCorp Vault: Secrets Management Fundamentals](https://mkabumattar.com/quizzes/post/vault-secrets-management-quiz), Master secrets management with this quiz on HashiCorp Vault engines, dynamic secrets, and the seal/unseal process.
- [Google Cloud Platform: GCP Essentials](https://mkabumattar.com/quizzes/post/gcp-essentials-quiz), Master the fundamentals of Google Cloud, including Project hierarchy, Compute Engine, GKE, and IAM.
- [Azure: Cloud Services & Architecture Fundamentals](https://mkabumattar.com/quizzes/post/azure-architecture-essentials-quiz), Master the fundamentals of Microsoft Azure, from core services like VMs and Networking to Azure DevOps and Resource Management.
- [Prometheus & Grafana: Monitoring & Observability Fundamentals](https://mkabumattar.com/quizzes/post/prometheus-grafana-monitoring-quiz), Test your monitoring and observability skills with this quiz on Prometheus metrics, PromQL, Grafana dashboards, and alerting logic.
- [Helm: Kubernetes Package Management Essentials](https://mkabumattar.com/quizzes/post/helm-kubernetes-package-management-quiz), Master Kubernetes package management with this quiz on Helm Charts, templates, values overrides, and release management.
- [GitLab CI/CD: Pipeline Automation Fundamentals](https://mkabumattar.com/quizzes/post/gitlab-cicd-pipeline-quiz), Test your mastery of GitLab CI/CD fundamentals, including .gitlab-ci.yml structure, Runner architecture, Pipeline stages, and environment management.
- [GitHub Actions: Workflow Automation Essentials](https://mkabumattar.com/quizzes/post/github-actions-workflow-automation-quiz), Master GitHub Actions fundamentals with this quiz covering Workflows, YAML syntax, the Actions Marketplace, Matrix builds, and secure secret management.
- [Jenkins: CI/CD Automation & Pipeline Management](https://mkabumattar.com/quizzes/post/jenkins-automation-fundamentals-quiz), Test your knowledge of Jenkins fundamentals with a quiz covering pipeline as code, Jenkinsfile syntax, agents, stages, build automation, plugins, distributed builds, Blue Ocean, and CI/CD best practices with Jenkins.
- [Terragrunt: Infrastructure as Code Management](https://mkabumattar.com/quizzes/post/terragrunt-iac-management-quiz), Test your knowledge of Terragrunt fundamentals with a quiz covering DRY principles, remote state management, module dependencies, configuration inheritance, run-all commands, hooks, and best practices for managing Terraform at scale.
- [CI/CD: Pipeline Automation Essentials](https://mkabumattar.com/quizzes/post/cicd-pipeline-essentials-quiz), Test your knowledge of CI/CD fundamentals with a quiz covering continuous integration, continuous delivery, pipelines, deployment strategies, automated testing, Infrastructure as Code, and DevOps best practices.
- [Git: Version Control Fundamentals](https://mkabumattar.com/quizzes/post/git-version-control-fundamentals-quiz), Test your knowledge of Git fundamentals with a quiz covering commits, branches, merging, rebasing, remote repositories, staging, conflict resolution, and version control best practices.
- [Kubernetes: Container Orchestration Essentials](https://mkabumattar.com/quizzes/post/kubernetes-orchestration-essentials-quiz), Test your knowledge of Kubernetes fundamentals with a quiz covering Pods, Services, Deployments, StatefulSets, storage, networking, RBAC, autoscaling, and container orchestration best practices.
- [TypeScript: Typed JavaScript Fundamentals](https://mkabumattar.com/quizzes/post/typescript-fundamentals-quiz), Test your knowledge of TypeScript fundamentals with a quiz covering static typing, interfaces, generics, type guards, enums, classes, advanced types, and TypeScript best practices.
- [JavaScript: Language Fundamentals & Modern Features](https://mkabumattar.com/quizzes/post/javascript-fundamentals-quiz), Test your knowledge of JavaScript fundamentals with this quiz covering variables, data types, functions, promises, DOM manipulation, ES6 features, closures, and modern JavaScript programming best practices.
- [Python: Programming Fundamentals & Best Practices](https://mkabumattar.com/quizzes/post/python-programming-fundamentals-quiz), Test your knowledge of Python fundamentals with this quiz covering data types, functions, loops, conditionals, list comprehensions, classes, exception handling, and Python programming best practices.
- [ArgoCD: GitOps Automation with Kubernetes](https://mkabumattar.com/quizzes/post/argocd-gitops-automation-quiz), Test your knowledge of ArgoCD fundamentals with this quiz covering GitOps principles, application deployment, sync policies, health status, ApplicationSets, and continuous delivery best practices for Kubernetes.
- [PowerShell: Windows Scripting Fundamentals](https://mkabumattar.com/quizzes/post/powershell-scripting-fundamentals-quiz), Test your knowledge of PowerShell fundamentals with a quiz covering cmdlets, pipelines, objects, variables, execution policies, functions, scripting best practices, and Windows automation.
- [Bash: Shell Scripting Fundamentals](https://mkabumattar.com/quizzes/post/bash-shell-scripting-fundamentals-quiz), Test your knowledge of Bash scripting fundamentals with a quiz covering variables, conditionals, loops, functions, I/O redirection, special parameters, and shell scripting best practices.
- [Linux: System Administration Fundamentals](https://mkabumattar.com/quizzes/post/linux-system-administration-fundamentals-quiz), Test your knowledge of Linux fundamentals with a quiz covering essential commands, file system navigation, permissions, process management, package managers, and Linux administration best practices.
- [Docker: Containerization Fundamentals](https://mkabumattar.com/quizzes/post/docker-containerization-fundamentals-quiz), Test your knowledge of Docker fundamentals with a quiz covering containers, images, Dockerfiles, networking, volumes, Docker Compose, and containerization best practices.
- [Ansible: Configuration Management & Automation](https://mkabumattar.com/quizzes/post/ansible-automation-fundamentals-quiz), Test your knowledge of Ansible fundamentals with a quiz covering playbooks, modules, inventory management, handlers, roles, and automation best practices.
- [AWS: Cloud Services & Architecture Fundamentals](https://mkabumattar.com/quizzes/post/aws-architecture-essentials-quiz), Test your knowledge of AWS fundamentals with a quiz covering key services, architecture, security, and best practices.
- [Terraform: Infrastructure as Code Essentials](https://mkabumattar.com/quizzes/post/terraform-essentials-quiz), Test your knowledge of Terraform fundamentals with this quiz covering key concepts and best practices.
# Roadmaps
8 entries.
- [Full-Stack Developer Beginner to Expert](https://mkabumattar.com/roadmaps/post/full-stack-developer-roadmap), A roadmap for learning full-stack development, covering frontend, backend, databases, APIs, DevOps, and deploying complete production applications.
- [Backend Developer Beginner to Expert](https://mkabumattar.com/roadmaps/post/backend-developer-roadmap), A roadmap for learning backend development, from a programming language and APIs to databases, caching, message queues, security, and scalable architectures.
- [Frontend Developer Beginner to Expert](https://mkabumattar.com/roadmaps/post/frontend-developer-roadmap), A roadmap for learning frontend development, from HTML, CSS, and JavaScript fundamentals to modern frameworks, state management, performance, and accessibility.
- [Release Engineer Beginner to Expert](https://mkabumattar.com/roadmaps/post/release-engineer-roadmap), A roadmap for learning Release Engineering, from version control and CI/CD fundamentals to advanced cloud automation, Infrastructure as Code, and GitOps delivery on AWS.
- [Site Reliability Engineer Beginner to Expert](https://mkabumattar.com/roadmaps/post/site-reliability-engineer-roadmap), A roadmap for learning Site Reliability Engineering, from Linux and networking fundamentals to advanced SLOs, observability, incident management, and automation on AWS.
- [Solutions Architect Beginner to Expert](https://mkabumattar.com/roadmaps/post/solutions-architect-roadmap), A roadmap for learning Solutions Architecture, from cloud fundamentals to advanced AWS design patterns, security, compliance, and scalable distributed systems.
- [DevOps Engineer Beginner to Expert](https://mkabumattar.com/roadmaps/post/devops-engineer-roadmap), A roadmap for learning DevOps engineering, from Linux fundamentals to advanced cloud-native and platform engineering concepts.
- [JavaScript Beginner to Expert](https://mkabumattar.com/roadmaps/post/javascript-developer-roadmap), A roadmap for learning JavaScript, from core language fundamentals to frontend, Node.js backend, and modern ecosystem tooling.
# Content Series
53 series.
- [Advanced Frontend Techniques](https://mkabumattar.com/series/advanced-frontend-techniques), 5 posts (Blog, Quizzes)
- [AI & LLM Engineering](https://mkabumattar.com/series/ai--llm-engineering), 7 posts (Blog, Codesnippets, Devtips)
- [Amazon Linux 2 LAMP Stack Setup](https://mkabumattar.com/series/amazon-linux-2-lamp-stack-setup), 4 posts (Blog)
- [APIs & Advanced Architecture](https://mkabumattar.com/series/apis--advanced-architecture), 8 posts (Blog, Quizzes)
- [AWS Automation](https://mkabumattar.com/series/aws-automation), 10 posts (Blog, Codesnippets)
- [AWS Bastion Host Setup](https://mkabumattar.com/series/aws-bastion-host-setup), 2 posts (Blog)
- [AWS CLI Guides](https://mkabumattar.com/series/aws-cli-guides), 5 posts (Blog, Cheatsheets)
- [AWS EC2 Interconnection](https://mkabumattar.com/series/aws-ec2-interconnection), 2 posts (Blog)
- [Backend Frameworks](https://mkabumattar.com/series/backend-frameworks), 1 post (Quizzes)
- [Backend Programming](https://mkabumattar.com/series/backend-programming), 7 posts (Blog, Quizzes)
- [Bash Snippets](https://mkabumattar.com/series/bash-snippets), 1 post (Codesnippets)
- [Career Roadmaps](https://mkabumattar.com/series/career-roadmaps), 8 posts (Roadmaps)
- [Certification Prep](https://mkabumattar.com/series/certification-prep), 9 posts (Flashcards)
- [CI/CD & GitOps](https://mkabumattar.com/series/cicd--gitops), 9 posts (Blog, Cheatsheets, Quizzes)
- [Cloud Platforms & Architecture](https://mkabumattar.com/series/cloud-platforms--architecture), 14 posts (Blog, Case studies, Glossary, Quizzes)
- [Config & Data Formats](https://mkabumattar.com/series/config--data-formats), 5 posts (Cheatsheets)
- [Container Security & DevSecOps](https://mkabumattar.com/series/container-security--devsecops), 7 posts (Blog, Devtips)
- [Containers & Kubernetes](https://mkabumattar.com/series/containers--kubernetes), 20 posts (Blog, Case studies, Cheatsheets, Devtips, Glossary, Quizzes)
- [Database Optimization](https://mkabumattar.com/series/database-optimization), 1 post (Codesnippets)
- [Databases & Data Persistence](https://mkabumattar.com/series/databases--data-persistence), 7 posts (Case studies, Cheatsheets, Quizzes)
- [Developer Environment & Tooling](https://mkabumattar.com/series/developer-environment--tooling), 10 posts (Blog, Cheatsheets)
- [DevOps & CI/CD Pipelines](https://mkabumattar.com/series/devops--cicd-pipelines), 12 posts (Blog, Devtips, Glossary)
- [DevOps & Infrastructure as Code](https://mkabumattar.com/series/devops--infrastructure-as-code), 5 posts (Glossary, Quizzes)
- [Docker Essentials](https://mkabumattar.com/series/docker-essentials), 4 posts (Blog, Flashcards)
- [FinOps & Cost Optimization](https://mkabumattar.com/series/finops--cost-optimization), 3 posts (Blog, Case studies)
- [Frontend Essentials](https://mkabumattar.com/series/frontend-essentials), 7 posts (Blog, Quizzes)
- [Frontend Frameworks & Mobile](https://mkabumattar.com/series/frontend-frameworks--mobile), 5 posts (Quizzes)
- [Git SSH Keys Setup](https://mkabumattar.com/series/git-ssh-keys-setup), 2 posts (Blog)
- [Infrastructure & Governance](https://mkabumattar.com/series/infrastructure--governance), 1 post (Devtips)
- [Infrastructure as Code Mastery](https://mkabumattar.com/series/infrastructure-as-code-mastery), 10 posts (Blog, Cheatsheets)
- [Jenkins CI/CD with AWS](https://mkabumattar.com/series/jenkins-cicd-with-aws), 3 posts (Blog)
- [Kubernetes & Container Orchestration](https://mkabumattar.com/series/kubernetes--container-orchestration), 1 post (Devtips)
- [Kubernetes Deep Dive](https://mkabumattar.com/series/kubernetes-deep-dive), 4 posts (Blog, Devtips)
- [Kubernetes Operations](https://mkabumattar.com/series/kubernetes-operations), 1 post (Devtips)
- [Linux & System Administration](https://mkabumattar.com/series/linux--system-administration), 19 posts (Cheatsheets, Codesnippets, Glossary, Quizzes)
- [Linux Essentials](https://mkabumattar.com/series/linux-essentials), 2 posts (Codesnippets, Devtips)
- [Mastering Terraform](https://mkabumattar.com/series/mastering-terraform), 11 posts (Blog, Cheatsheets, Devtips)
- [Monitoring, Security & Infrastructure](https://mkabumattar.com/series/monitoring-security--infrastructure), 5 posts (Quizzes)
- [Node.js Backend Essentials](https://mkabumattar.com/series/nodejs-backend-essentials), 1 post (Codesnippets)
- [Node.js Express TypeScript Setup](https://mkabumattar.com/series/nodejs-express-typescript-setup), 2 posts (Blog)
- [Node.js Snippets](https://mkabumattar.com/series/nodejs-snippets), 1 post (Codesnippets)
- [Observability & Monitoring](https://mkabumattar.com/series/observability--monitoring), 5 posts (Blog, Devtips)
- [Performance & Scaling](https://mkabumattar.com/series/performance--scaling), 1 post (Codesnippets)
- [Platform Engineering](https://mkabumattar.com/series/platform-engineering), 3 posts (Blog, Case studies)
- [Programming Languages](https://mkabumattar.com/series/programming-languages), 12 posts (Blog, Cheatsheets, Codesnippets, Flashcards, Quizzes)
- [Python Snippets](https://mkabumattar.com/series/python-snippets), 1 post (Codesnippets)
- [QuenchWorks](https://mkabumattar.com/series/quenchworks), 1 post (Case studies)
- [Reliability & Resilience Engineering](https://mkabumattar.com/series/reliability--resilience-engineering), 2 posts (Blog)
- [Secret Management](https://mkabumattar.com/series/secret-management), 1 post (Codesnippets)
- [Shell & Terminal Productivity](https://mkabumattar.com/series/shell--terminal-productivity), 1 post (Codesnippets)
- [Shell Mastery](https://mkabumattar.com/series/shell-mastery), 2 posts (Codesnippets)
- [Software Engineering Craft](https://mkabumattar.com/series/software-engineering-craft), 7 posts (Blog, Quizzes)
- [Spring Boot on AWS](https://mkabumattar.com/series/spring-boot-on-aws), 2 posts (Blog)