If releasing a new version means a maintenance window, releases get postponed, grouped into large batches and done late at night. Large, rare releases are riskier than small, frequent ones. Being able to deploy at any time without users noticing changes how a team works, and it is achievable with modest infrastructure.
Where downtime comes from
There are three usual causes. The old version is stopped before the new one is ready to serve requests. Requests in progress are cut off when the old process is killed. Or a database change is applied that the running code cannot work with. The first two are solved by the deployment method. The third needs a discipline of its own and causes most of the real trouble.
Rolling deployment
With several copies of the application running behind a load balancer, replace them one at a time. Take one copy out of rotation, update it, check that it is healthy and return it. Repeat for the rest.
This needs little extra capacity and is simple to automate. During the rollout, old and new versions serve requests side by side, so the two must be compatible with each other. Rolling back means repeating the process in reverse.
Blue-green deployment
Keep two complete environments. One, called blue, serves all traffic. Deploy the new version to the other, green, test it there, then switch traffic across in one step. If something is wrong, switch back.
The advantages are a clean cut-over and a near-instant rollback. The costs are running double capacity during the release and the fact that both environments normally share one database, which brings back the compatibility question.
Canary release
Send a small share of traffic, perhaps five percent, to the new version and watch its error rate and response times. If they look normal, increase the share in steps until it takes everything. A problem then affects few users and is caught early.
This requires routing that can split traffic by percentage and monitoring good enough to compare the two versions. It suits products with enough traffic for five percent to be a meaningful sample.
The database is the hard part
In all three methods, old and new code run against the same database for a while. A migration that renames a column breaks the old code the moment it is applied. A migration that adds a required column breaks old code that does not fill it.
The technique that solves this is called expand and contract. Every schema change is split into steps that are each compatible with the code running at that moment.
- Expand. Add the new column or table, leaving the old one in place. Old code is unaffected.
- Migrate. Deploy code that writes to both and reads from the new one. Copy existing data across in the background.
- Contract. Once nothing uses the old column, remove it in a later release.
A rename therefore takes two or three releases instead of one. That is the price of never being offline, and it also means every step can be rolled back.
Keep migrations short
Some schema changes lock a table while they run, and on a large table that can block the application for minutes. Add indexes with the option that builds them without blocking writes. Change large amounts of data in small batches, not in a single statement. Test migrations against a copy of production data, since a change that is instant on a small table can take an hour on the real one.
Health checks and shutdown
The load balancer should send traffic to a new copy only when the application reports that it is ready, meaning it has started and can reach its database. When stopping an old copy, first remove it from rotation, let requests in progress finish, then stop it. Without this, every deployment drops a handful of requests.
Background workers
Queued jobs created by the old version may be picked up by the new one, and the reverse. Keep job formats compatible across one release, in the same way as the database schema.
Plan the rollback
Know before each release how to return to the previous version and whether that is still safe after the migration has run. With expand and contract it is, because the old structure is still present. Destructive changes are held back until the new version has proved itself.
Feature flags
A feature flag lets code be deployed switched off and turned on later for some or all users. This separates the act of deploying from the act of releasing. A risky feature can be turned off in seconds without a new deployment.
How much do you need?
For an internal tool used in office hours, a thirty-second restart in the evening may be perfectly acceptable. For a customer-facing product, a rolling deployment with health checks and careful migrations covers most needs. Blue-green and canary releases are worth their extra cost when traffic and risk are higher. Choose the simplest method that meets the real requirement.
Summary
Start new copies before stopping old ones, let requests finish, and split every database change into compatible steps. Pick rolling, blue-green or canary according to your traffic and tolerance for risk. Deployment pipelines of this kind are part of our cloud and DevOps service. If your releases still need a maintenance window, we can help remove it.