SREs split into two camps, and neither one is fully right.
The first camp matured out of operational IT. They’ve lived through a bad patch Tuesday, a driver update that bricked a fleet, a vendor hotfix that took down a cluster. Their instinct: let a patch cycle pass before you touch it. Let someone else find the landmine. This dates me, but is the world I learned in myself.
The second camp never did operational IT work. They came up through cloud-native, IaC, CI/CD. They don’t have the scar tissue, but they inherit the doctrine anyway, because the first camp is louder and… has tenure.
Both camps are operating on a rule that was correct for a world that doesn’t exist anymore.
Where the doctrine came from
“Wait a cycle” wasn’t stupid when it was invented. It was a rational response to a specific threat model: systems sitting comfortably behind a firewall, behind NAT, with a slow-moving, mostly-trusted internal network. In that world, the cost of a bad patch (downtime, a broken driver, an angry VP) outweighed the cost of a slow patch, because the exposure window barely mattered. Nothing outside the perimeter could reach the vulnerable system anyway.
That threat model is gone. IoT, always-connected infrastructure, cloud workloads with public ingress, supply-chain attacks that don’t care about your patch calendar. The exposure window is no longer a rounding error. It’s the entire risk.
The math flipped
Run the actual comparison: probability and blast radius of a bad patch, versus probability and blast radius of running a known-exploitable version for an extra two to four weeks because “that’s the cycle.”
In a perimeter-protected, low-connectivity environment, delay was the safer bet. In a heavily connected, internet-facing, IoT-laden environment, delay is frequently the worse bet. The scales tipped and most of the doctrine’s adherents never recalculated.
The humility problem
The old-hat camp will tell you they’re being careful. What they’re actually doing is claiming they can manually vet every change in a modern software stack before it ships to prod. That’s not caution, that’s a false claim of competence. Nobody is reading millions of lines of interconnected upstream diffs across a dependency graph before greenlighting a patch. The manual gate people think they’re running doesn’t exist in the form they think it does. The software stack you’re reading this post on today is millions of lines of code: the OS, the driver layers, the browser, the software running on top of all that like accessibility, plugins, etc.
The second camp goes along with it because they don’t have the operational scar tissue to push back, and the first camp has institutional authority, after all, nobody likes dealing with broken software.
You can do both
This isn’t an argument for reckless auto-patching with no controls. It’s an argument that “patch immediately” and “have eyes on the system” aren’t mutually exclusive.
Automate patching. Also build alerting robust enough to catch it fast when an automated patch breaks something. Make sure you’re patching pre-production systems first, and make sure your alerting on those is sufficient to flag problematic releases verses just assuming it. You cannot manually patch at the speed the current threat landscape requires. Manual review as your primary gate is already too slow the day you adopt it. The gate has to be automated deployment plus fast detection, not automated deployment blocked by manual approval.
Semver will not save you
The industry treats semver like a contract. It isn’t reliably one, because vendors and especially OSS maintainers don’t follow it consistently. Breaking changes routinely land in patch and minor releases. This isn’t rare, it’s routine enough that “it’s a patch version” tells you almost nothing about blast radius.
There’s a second failure mode that’s worse: not all patches get backported. Say you’re pinned like this:
"some-package": "^1.3.0"
That constraint reads like safety. It isn’t, if the vendor has already moved to 2.0 and stopped backporting fixes to the 1.x line. You’re sitting on 1.3, believing you’re getting patched, and you’re actually flying blind. No new CVE fixes are coming to your branch. The version constraint gave you false confidence, not protection.
The only thing that catches this is manual, eyes-on inspection of your dependency tree over time; not once at adoption, on an ongoing basis. Not some vendor promise of AI slop. Checking whether the upstream project is still maintaining the line you’re pinned to. Semver compliance is not something you can assume, it’s something you have to verify.
Version pinning isn’t absolute either
Pinning is presented as the fix for semver’s unreliability. It helps, but it’s not airtight. Some dependencies have overwritten their own published version pins after the fact; republishing under a version number that’s already shipped, with different contents. Not best practice, not supposed to happen, happens anyway.
The only pinning that actually guarantees you’re running what you think you’re running is pinning to a content hash, not a version string. Almost nobody does this in practice, because it’s operationally painful; every hash has to be tracked and updated manually with each intentional bump, and most tooling doesn’t make it frictionless. So most orgs run with a security model they believe is stronger than it is.
Where this leaves you
None of this argues for going back to manual gatekeeping. It argues for being honest about what your controls actually give you:
- Automated patching, not manual sign-off, is the only thing fast enough to matter against current exposure windows.
- Alerting has to be good enough to catch a bad automated patch quickly, because you’re trading manual review for speed and need to cover that gap somewhere.
- Semver compliance is not guaranteed. Check whether your pinned line is still receiving backports, not just whether your constraint syntax looks safe.
- Version pinning reduces risk, it doesn’t eliminate it. Hash pinning is the only real guarantee, and most teams won’t pay the operational cost to do it.
- “Wait a cycle” as a blanket policy is a legacy heuristic from a threat model that no longer applies to most connected systems. It needs to be replaced with actual risk math, not tenure-based instinct.