Skip to main content

Stability checklist

EveryoneStandardUpdated October 1, 2026

A Stability level says in one word how stable something is. This checklist shows what's behind that word: for each area, such as tests or security, what should be in place at each level. Clients and engineers read the same list, so both sides expect the same thing.

The checklist is what we aim for, not a gate. Every item checked is the ideal. When an item applies and isn't checked, we say so, so you can judge the risk.

Prove it before you harden it

Moving through the levels keeps us from spending time hardening something before we know it solves the right problem, in the right way. That's why most of the checklist arrives at Beta and GA.

  • The client sees and tries the work at every level. As a rule of thumb, at least once per level, and usually several times as it changes.
  • Don't go past Beta on a guess. Before GA, we should be highly confident, from the client's own use, that it solves the right problem in the right way.

Frequent client feedback is the key. It tells us when to move up, when to change direction, and when to stop. Moving between levels has an example.

The checklist​

Choose how to view it, then show everything or pick one:

  • By category shows each category, such as Safe, with what its sub categories need at each level.
  • By stability level shows everything that should be in place at a level, such as Beta.
  • By sub category shows a sub category, such as Tests, at each level.

Each level's list is complete on its own, including what carries over from the levels below. Hover over a sub category for why it matters.

The security baseline doesn't change with the level

The security and privacy baseline, and the commercial rules, hold at every level, which is why the security items start at Prototype. GA adds reviews and hardening on top.

PrototypeExploratory

Works correctly

Tests
  • Demo path checked by hand.
Edge cases and errors
  • Known gaps written down.
Data and migrations
  • Sample data clearly marked.
Performance
  • Fast enough to demo.

Safe

Security and privacy
  • No secrets or real client data in code, mocks, or demos.
  • Auth enforced on anything touching client systems or data.
  • Secrets kept in the secret store.
  • Personal data collected only where needed.
  • Protected against injection and cross-site scripting.
Dependencies and integrationsNothing required yet.

Runs in production

Logs, traces, and metricsNothing required yet.
Monitoring and alertingNothing required yet.
Release and rollback
  • Outside production, or behind a flag that's off.
Support and fixesNothing required yet.

Usable and lasting

User experience
  • Shows the idea clearly enough to react to.
Documentation
  • A note on what it shows, what's mocked, and how to run it.
Code and maintainability
  • Kept apart from production code.
Breaking changesNothing required yet.

AlphaUnstable

Works correctly

Tests
  • Automated test for the main path.
Edge cases and errors
  • Known gaps written down.
  • Errors fail visibly, never silently.
  • Input validated where it enters.
Data and migrations
  • Sample data clearly marked.
  • Schema changes are migrations.
  • Data it creates can be cleaned up.
Performance
  • Usable at today's scale.

Safe

Security and privacy
  • No secrets or real client data in code, mocks, or demos.
  • Auth enforced on anything touching client systems or data.
  • Secrets kept in the secret store.
  • Personal data collected only where needed.
  • Protected against injection and cross-site scripting.
Dependencies and integrations
  • Dependencies maintained, with compatible licenses.
  • Service URLs and keys in config, not code.

Runs in production

Logs, traces, and metrics
  • Errors logged with context, without secrets or personal data.
  • Logs sent to the project's log store.
Monitoring and alertingNothing required yet.
Release and rollback
  • Behind a feature flag, with an Alpha badge.
Support and fixes
  • Bugs logged as issues.
  • Client knows how to report a problem.

Usable and lasting

User experience
  • A real user can complete the main workflow.
Documentation
  • Users: how to try it, and what's missing.
  • Developers: how to run it, and its configuration.
Code and maintainability
  • Follows conventions, and passes lint and type checks.
  • Reviewed by another engineer.
Breaking changes
  • Breaking changes noted in the pull request.

BetaApproaching stability

Works correctly

Tests
  • Automated test for the main path.
  • Edge cases and error paths tested.
  • Tests run in CI and block merging.
Edge cases and errors
  • Known gaps written down.
  • Errors fail visibly, never silently.
  • Input validated where it enters.
  • Empty, missing, duplicate, and large inputs handled.
  • Helpful error messages, not stack traces.
Data and migrations
  • Sample data clearly marked.
  • Schema changes are migrations.
  • Data it creates can be cleaned up.
  • Data validated before it's written.
  • Migrations reversible, or backed up.
Performance
  • Usable at today's scale.
  • Tried at realistic volumes, with bounded queries.
  • Obvious problems fixed, such as N+1 queries and missing indexes.

Safe

Security and privacy
  • No secrets or real client data in code, mocks, or demos.
  • Auth enforced on anything touching client systems or data.
  • Secrets kept in the secret store.
  • Personal data collected only where needed.
  • Protected against injection and cross-site scripting.
Dependencies and integrations
  • Dependencies maintained, with compatible licenses.
  • Service URLs and keys in config, not code.
  • Versions pinned and scanned for vulnerabilities.
  • Timeouts on external calls, and retries where safe.
  • Fails clearly when an external service is down.

Runs in production

Logs, traces, and metrics
  • Errors logged with context, without secrets or personal data.
  • Logs sent to the project's log store.
  • Structured logs with a request ID.
  • Errors reported to the error tracker.
  • Requests traced across services.
  • Rate, error, and duration metrics for key endpoints.
Monitoring and alerting
  • Health checks, watched by uptime monitoring.
  • Dashboard for key metrics.
Release and rollback
  • Deployed through the normal pipeline.
  • Beta badge wherever people see it.
  • Can be turned off or rolled back without a code change.
Support and fixes
  • Bugs logged as issues.
  • Client knows how to report a problem.
  • Bugs triaged and fixed as found.
  • Known issues shared with the client.

Usable and lasting

User experience
  • A real user can complete the main workflow.
  • Loading, empty, and error states designed.
  • Keyboard accessible, with labels and readable contrast.
Documentation
  • Developers: how to run it, and its configuration.
  • Users: how to use and configure it.
  • Developers: architecture overview and key decisions.
  • User-facing changes in release notes.
Code and maintainability
  • Follows conventions, and passes lint and type checks.
  • Reviewed by another engineer.
  • Business logic separate from UI, storage, and services.
  • Earlier shortcuts removed or tracked.
Breaking changes
  • Breaking changes noted in the pull request.
  • Breaking changes avoided, or announced with migration notes.

GAStable

Works correctly

Tests
  • Automated test for the main path.
  • Edge cases and error paths tested.
  • Tests run in CI and block merging.
  • End-to-end tests for critical workflows.
  • Regression test for every fixed bug.
  • No flaky or skipped tests without an issue.
Edge cases and errors
  • Known gaps written down.
  • Errors fail visibly, never silently.
  • Input validated where it enters.
  • Empty, missing, duplicate, and large inputs handled.
  • Helpful error messages, not stack traces.
  • Concurrency and partial failures handled, and retries are safe.
  • Every known edge case handled or documented as out of scope.
Data and migrations
  • Sample data clearly marked.
  • Schema changes are migrations.
  • Data it creates can be cleaned up.
  • Data validated before it's written.
  • Migrations reversible, or backed up.
  • Migrations tested on production-sized data.
  • Backups with a tested restore.
Performance
  • Usable at today's scale.
  • Tried at realistic volumes, with bounded queries.
  • Obvious problems fixed, such as N+1 queries and missing indexes.
  • Meets performance targets, measured.
  • Handles peak load with headroom.

Safe

Security and privacy
  • No secrets or real client data in code, mocks, or demos.
  • Auth enforced on anything touching client systems or data.
  • Secrets kept in the secret store.
  • Personal data collected only where needed.
  • Protected against injection and cross-site scripting.
  • Security review for sign-in, payments, and personal data.
  • Least privilege, and an audit trail for sensitive actions.
  • Personal data can be exported and deleted on request.
Dependencies and integrations
  • Dependencies maintained, with compatible licenses.
  • Service URLs and keys in config, not code.
  • Versions pinned and scanned for vulnerabilities.
  • Timeouts on external calls, and retries where safe.
  • Fails clearly when an external service is down.
  • Depends on nothing less stable than GA.
  • Rate limits, quotas, and costs fit expected use.
  • Dependency updates automated or scheduled.

Runs in production

Logs, traces, and metrics
  • Errors logged with context, without secrets or personal data.
  • Logs sent to the project's log store.
  • Structured logs with a request ID.
  • Errors reported to the error tracker.
  • Requests traced across services.
  • Rate, error, and duration metrics for key endpoints.
  • Workflow and business metrics.
  • End-to-end traces for critical workflows, linked from logs.
  • Log retention fits privacy needs.
Monitoring and alerting
  • Health checks, watched by uptime monitoring.
  • Dashboard for key metrics.
  • Latency, error rate, and availability targets on the dashboard.
  • Alerts on user-facing failures, routed to whoever operates production.
  • Every alert has a runbook note.
  • Alerts test-fired, and tuned to avoid noise.
Release and rollback
  • Deployed through the normal pipeline.
  • Badge and feature flag removed.
  • Rollback tested.
Support and fixes
  • Bugs logged as issues.
  • Client knows how to report a problem.
  • Bugs triaged and fixed as found.
  • Known issues shared with the client.
  • GA bugs fixed before new work.
  • Client knows who to contact, and what response to expect.

Usable and lasting

User experience
  • A real user can complete the main workflow.
  • Loading, empty, and error states designed.
  • Keyboard accessible, with labels and readable contrast.
  • Consistent with the product on supported browsers and devices.
  • Meets the project's accessibility standard.
Documentation
  • Developers: how to run it, and its configuration.
  • Developers: architecture overview and key decisions.
  • User-facing changes in release notes.
  • Users: complete, current documentation.
  • Developers: reference docs for public APIs.
  • Operations: runbook to deploy, roll back, and diagnose.
  • Changelog entry for the release.
Code and maintainability
  • Follows conventions, and passes lint and type checks.
  • Reviewed by another engineer.
  • Business logic separate from UI, storage, and services.
  • Earlier shortcuts removed or tracked.
  • Another engineer can change it without the author.
  • No untracked TODOs or dead code.
Breaking changes
  • Breaking changes noted in the pull request.
  • Public interfaces documented as stable.
  • Breaking changes only in a new version.

To change an item, propose an edit to src/data/stability-checklist.ts. The checklist on this page is generated from it.

Frequently asked questions​

Does every item have to be checked before something is called a level?

No. Checking every item is what we aim for, not a requirement. When an item that applies isn't checked, the engineer says so. The exception is the security and privacy baseline, which holds at every level.

What does Beta not include?

View the checklist by stability level and compare Beta with GA. In short: alerts, end-to-end traces, tested restores, a runbook, and the promise that breaking changes only come in a new version. Those come with GA.

Why does a GA feature need GA dependencies?

A feature is only as stable as what it relies on. A GA report built on a beta API can break when that API changes, however carefully the report was built.

What do traces give us that logs don't?

Logs say what happened in one place. A trace follows one request through every service, query, and external call, with how long each step took, so it shows where something slow or broken went wrong.

Who operates the alerts a GA feature needs?

Whoever operates production. On Exalynt Managed, that's Exalynt. On Self Managed, alerts go to the client's team.