Make a staging OTP shortcut fail at startup, not at verification
A fixed code that opens every account is the fastest way to test a login flow, and a breach on the day it reaches production. The awkward part is that the verification function cannot catch this: by the time a candidate string arrives, the only question left is whether it matches, and a universal code matches. The defence has to sit earlier, in the code that reads the environment and in the code that wires the application together.
We build and run authentication flows for systems where a login is the door to money, so we had to decide where a convenience like this is allowed to live. Our answer is that it may exist, but only as an object you can interrogate before the first request is served. We maintain a small library for this, otpguard, and the test-mode module is the part we argued about most.
The check is the wrong place to put the question#
The usual shape of this bug is a conditional inside the verification path:
if candidate == settings.OTP_TEST_CODE and settings.DEBUG:
return TrueTwo separate things go wrong here, and only one of them is about the comparison. The first is that an environment question, is this deployment allowed to have a bypass, gets answered per request, inside an authentication decision, with nobody watching. The second is that the answer leans on a flag that is routinely wrong in exactly the deployments where it matters: a staging-like box promoted to serve real traffic, a container that inherited an env file, a DEBUG that defaults to true in a path nobody audited.
So we moved the question to two moments that happen once, loudly, and before any user is involved: when the mode is constructed from the environment, and when the application assembles its dependencies.
Constructing the mode is where the environment gets interrogated#
TestMode is off unless you pass a code, and you cannot pass a code without a lifetime:
from datetime import timedelta
from otpguard import TestMode
TestMode() # off
TestMode("424242") # ValueError: test mode must expire
TestMode("424242", ttl=timedelta(days=7)) # ValueError: ttl must not exceed 1 day
TestMode("424242", ttl=timedelta(hours=1)) # armed, and logs at CRITICALMAX_TEST_MODE_TTL is twenty-four hours, and it is a ceiling rather than a default. Whatever the caller asks for, a universal code cannot outlive a working day, so a bypass left in an environment file stops being a credential on its own, without anyone remembering it. Passing a ttl with no code is an error too, not a no-op, because a half-removed configuration is something we would rather hear about than quietly honour.
Reading the same settings from the environment goes through from_env, which is where the deployment gets to disqualify itself:
import os
from otpguard import (
ENVIRONMENT_VAR,
TEST_MODE_CODE_VAR,
TEST_MODE_TTL_VAR,
TestMode,
)
TestMode.from_env({}) # off
TestMode.from_env({TEST_MODE_CODE_VAR: " "}) # off: a blank code is no code
env = {
TEST_MODE_CODE_VAR: "424242",
TEST_MODE_TTL_VAR: "600",
ENVIRONMENT_VAR: "Production",
}
TestMode.from_env(env) # TestModeInProduction
mode = TestMode.from_env(os.environ) # the real call sitePRODUCTION_ENVIRONMENTS holds live, prod and production. That list is deliberately small and dumb. It is not a security boundary, since an environment can call itself anything, but it does catch the most common version of the accident, which is a value copied between env files along with everything else. A deployment that announces itself as production and also carries a universal code has told you two contradictory things, and the library refuses to settle that contradiction in favour of convenience.
The guard asks whether it is configured, not whether it works#
The second moment is wiring, and the function there is require_no_test_mode. Its rule is narrower than it first looks: it raises while a universal code is configured at all, expired or not.
import pytest
from datetime import timedelta
from otpguard import TestMode, TestModeInProduction, require_no_test_mode
def test_a_configured_mode_is_refused_where_it_must_not_exist():
with pytest.raises(TestModeInProduction):
require_no_test_mode(TestMode("424242", ttl=timedelta(hours=1)))
def test_a_mode_that_is_off_passes_the_guard():
mode = TestMode()
assert require_no_test_mode(mode) is modeLetting an expired mode through the guard would be the wrong behaviour, and the reason is operational rather than cryptographic. An expired universal code is still a value sitting in a deployment's configuration, still travelling through whatever pipeline put it there, still one TTL bump or one clock argument away from being live again. is_active() is a question about now, while is_configured is a question about the deployment. Startup cares about the second one.
That makes the wiring code boring, which is what we want from it:
import os
from otpguard import TestMode, require_no_test_mode, require_real_sender
def build_code_flow(settings, sender):
mode = TestMode.from_env(os.environ)
if settings.is_production:
require_no_test_mode(mode)
require_real_sender(sender)
return mode, senderNote that settings.is_production here is the application's notion of production, not the library's. The two checks are redundant on purpose: from_env refuses a deployment that names itself production, and the guard refuses a mode however it was built, including one constructed in code, in a test fixture that escaped, or under an environment name the library has never heard of.
The same shape for anything that only pretends#
A universal code is one member of a family: development objects that accept real input and produce no real effect. The other one in this library is StubSender, which keeps messages instead of delivering them so that local runs and tests can read the code back. Left in a deployment it accepts every code and delivers none of them, which is the same failure wearing different clothes, an auth flow that appears to work and verifies nothing.
So it carries the same two properties. It admits what it is with a is_stub = True marker that callers can ask about without knowing the concrete type, and there is a guard with the same signature:
from otpguard import StubSender, StubSenderInProduction, require_real_sender
require_real_sender(StubSender())
# StubSenderInProduction: StubSender delivers nothing; configure a real senderWhen the two guards share one shape, you can read the startup block above as a single policy, this deployment uses nothing that pretends, rather than as two unrelated checks someone might implement only half of.
Three log levels, and a repr that keeps the secret#
Arming the mode logs at CRITICAL, naming the moment it expires. Every accepted universal code logs at ERROR, stating that no stored code was verified. A code offered after expiry is refused and logged at WARNING. Those levels are not for the developer who turned it on. They are for whoever is grepping a production log months later and needs the arming event, the uses, and the probes to be three distinguishable things.
Two details make the logs trustworthy. First, verification tries the stored digest before it considers the universal code, so a real code produces no test-mode log line at all and the bypass cannot shadow a genuine one:
mode.verify(candidate, stored, pepper=SERVER_SECRET)Second, the secret stays out of everything that gets printed. repr(TestMode()) is TestMode(off), an armed mode's repr does not contain the universal code, and Message renders its code as '***', because a live code in a traceback is a code someone else can use.
What it costs to run#
Less than the conditional it replaces, with one honest exception. At startup you read the environment once and build an object. At verification you pass that object instead of calling a module-level function. Nothing is added to the request path, and the verification call stays unconditional, which is the main win: there is no if DEBUG branch left to be wrong.
The cost is that staging has to re-arm at most daily, because the TTL ceiling will not let a bypass sit there for a sprint. We consider that the feature. The other cost is real and worth planning for, because a deployment that carries a universal code into production now refuses to start. That has to show up in the pipeline as a failed release with a readable message, not as a container restarting in a loop that nobody is watching. The exception text is written for that moment, test mode must not be configured here followed by the reason, since the person reading it is probably mid-incident and did not write the config.
A bypass you can ask about, that expires on its own, announces itself three ways and can stop the process, is still a bypass. It is no longer a credential with no owner and no end date. The convenience is allowed to exist. Existing quietly is the part that is not.