It's OK to hardcode feature flags (2025)(code.mendhak.com) |
It's OK to hardcode feature flags (2025)(code.mendhak.com) |
The config approach is critical for using feature flags as a software development lifecycle tool. It is how you manage having a codebase which contains the unfinished code for new feature x, but can still be deployed and pass all tests without feature x being turned on.
In this model you need a mechanism which allows a developer who is working on feature x to enable it for local testing, and for your CI system to be able to interact with the flag system to test that the application works in both states - with x turned on and off.
This is ideal for trunk based development models; feature branches are an alternative approach that doesn’t really benefit from this (indeed it adds complexity to working in feature branches).
Meanwhile feature flag services are to solve the problem that different people using the same software need different features turned on. That can be as simple as internal testers or beta users, it can be holding features to roll out in fixed update windows per tenant, or it can be part of a risk management strategy where features are rolled out through progressive exposure.
It can also be tempting to mix up your feature flags system with an A/B testing system - you can use a feature flag service to expose a feature to a test cohort and measure performance changes.
There can be reasons for doing that but it’s really important not to tie all these things together: not every development lifecycle change is an ab test hypothesis. Not every ab test hypothesis is a development lifecycle change (often it is really about testing changes in data, and feature config is just one piece of data you might want to change). Similarly some other data changes than code changes need to roll out progressively to mitigate risk.
So all these things might be feature-flag shaped, but that doesn’t mean you can substitute different feature flag solutions in and solve the same problems.
This post is saying ‘it’s okay not to have an exposure control solution; you can have a config file’ - which is obviously true, if what you have is a config management problem, not an exposure control problem.
For Quonfig I landed on:
- Data model wise they are identical. Flags and configs can both be targeted. They can both use segments. They can both do partial rollouts. They can both have the same range of values (bool/number/duration/json/json-w-schema/etc)
- Its fine to use Flags for your Experiments/AB Tests, but that's just the "allocation engine". The rest of experimentation is the exposure tracking and goal tracking. Those ought live in your product analytics stack, because they are really just events. But its helpful to have a single place (flags) for the experiment allocation because then you can re-use segments for things like hold out groups / and just general targeting.
- The only real difference is that Flags are intended to be ephemeral and configs are intended to be permanent. A good UI should show you how long a flag has been alive and help you clean it up (or convert it to a config) if it's been true for everyone for too long.
- The use cases are different enough that it's worth keeping them in two separate UI.
Examples:
Experiment flags for A/B testing
Permission checks at a user level
Product flags for feature gating by plan
Rollout flags to launch a new feature gradually
Configuration for an account/system/feature
The number of times I had to push against the “just put it behind hasFeatureFlag(user, flag)” is more than I can count at this point! I think it comes from a misunderstanding of the DRY principle, to be honest.
entitlements is a huge one. Sooooo many orgs are in an absolute rats nest of confusion because half their entitlements are in something like an auth system, half are in something like billing and the third half is hacked into feature flags.
One particular pain point is that finance or something "do users with access to feature X retain better and it's super hard to figure out.
My overall feel is that developers are trying to solve a problem and that it turns out the configuration is as fundamentally as important as the software itself. I define "dynamic configuration" as "a key value store which takes context and has rules based system to give you a value". This primitive turns out to be extraordinarily useful and powerful. Rather than try to split it into N different systems, each custom / don't have telemetry / aren't available to all services SDKs. What if you have ONE BIG CONFIG system that really does a bang up job.
(Note: copying the big boys is not always the right move, but https://research.facebook.com/publications/holistic-configur... does give some credence that this is a decent approach)
By focussing on one terrific system, you can put all your eggs into the basket of making that reliable / comprehensible / flexible / strict.
The one on your list that I think DOESN'T belong would be permission checks at user level. The N of that is not a good fit for config (though when you look at the facebook paper, it's pretty wild how they've scaled it (but it doesn't do permissions afaik)).
Every project should have flags, but many projects need just the basics and a service is overkill.
Rolling your own JSON still feels like something we ought avoid though. Yes to start out it’s 95% booleans. But then you want a rollout. And then you want some targeting rules. And then you want non booleans… maybe some json. Oo wouldn’t it be nice if the json could conform to a schema… and then eventually you are like damn I really want to change these without deploying. Or you want to read the same flag from multiple services.
I’ve tried to incorporate this lowest common denominator into https://quonfig.com Use it totally free & open source as SDK, and it’s just loading JSON that you can track in git. Agents love it, hot reloads, SDK in lots of languages. But vs rolling your own you’ve got a lot of headroom on the design. A bunch of targeting operators. Segments etc. And then if you do want to get a nice UI / delivery network for real time updates, then you can use the paid side of things.
Local use description: https://docs.quonfig.com/docs/how-tos/open-source-local
something like enabledFeature = (flag1 || flag2 || flag3) && flag4
then, down the road, you remove flag4. Hopefully the above would result in a compile error but you may be in a language or situation where a missing flag4 is interpreted as boolean false. That could cause all kinds of havoc in logical combinations like that are scattered all over the codebase. Even worse would be REST APIs retrieving flag values because who knows what's happening in that service code? Plus, you'd never know there was an issue until users start reporting missing or extra features showing up unless you have e2e tests for every possible combination of feature flags...
Immediately create story to remove said feature flag controlling access to the feature and review it during backlog refinements.
In my experience, engineers aren't using them to account for managerial dithering, they're doing it for safe deployments and experiments and rollouts and such. A product with millions of users can easily have a tens or even hundreds of active switches at any moment (I'm assuming a large engineering team behind said product), and that's not necessarily a bad thing.
However, as someone else noted here, you absolutely MUST delete and clean up your flags/gates/whatever when you've completed that effort. That part can be tricky because not everyone has the discipline to pay off tech debt.
Usually, a flag/gate should not live in code for more than a few months. If it does, it should have robust justification.
1. For reducing outliving, I always create a ticket to remove the flag at the same time it gets added in. This way there's documented work that will get scheduled. YMMV depending on how your team does planning.
2. For flag branch implementations, it's often fine to do an if/else, but I've seen this blow up into a real headache. When I can, I like having two implementations of an interface that get swapped between. Code calls the interface as normal the the underlying implementation is the same. Works for React components and larger implementations/changes that already have an interface. Don't want to force an interface where it feels wrong.
There are no solutions, only tradeoffs.
I'll give some examples:
Firm 1 had a single server running tron [0] for ALL scheduling of processes.
Pros: very easy to see what should run when and made audits a breeze
Cons: the central server died and it was a giant outage response to make sure that when the server came up it didn't start killing processes that should be running
Firm 2 used a management gui to create custom cron entries on each machine
Pros: each node had a local copy of the schedule and could keep going even if the central system died
Cons: each node had a local copy of the schedule which could "drift" from other nodes, a node could be forgotten etc
So, more generally, I agree that it's good to label things as "ok" so that we don't get into flame wars etc. That being said, the more important point is to say that if you are going to pick a strategy, do the work to support, build tooling and plan for outages related to that strategy.
Specifically, the idea of putting toggles on risky functionality is a good idea that is useful broadly. Being able to dynamically toggle feature flags without a restart, or apply some features to some segment of users is unlikely to be useful and entails a lot of complexity and potential footguns.
My approach to feature flags is even simpler than what is suggested here: feature flags are just a section in your config. Done. Even if you are confident that you need more advanced capabilities, you should not "advance" to that level until you have demonstrable proof that your team can manage simple config values for it.
When something does go wrong, the compiler degrades gracefully, and still can make sense of most of the code.
MS gets a lot of flak (sometimes deservedly so), but they can sometimes also show what a large organization of skilled individuals working under consistent direction is capable of.
#ifdef FEATURE_AI_V2
;)I can't hard code any of the flags until I have big customers notified and onboarded with the change.
So one would say why aren't your customers onboarded with the change before you implement it. Well no because they have 20-30 other applications they are using on daily basis.
While yeah we can push some customers to new features, we still need to have some realistic adoption in place before we can switch ... mostly because we are not Attlasian or Microsoft — we have actual people taking angry calls and angry e-mails from the customers.
Any configuration read out of a JSON file is not hardcoded. Hardcoded means you need to recompile to change it.
(And yes, hardcoded, as in flags set with #define or equivalent, are totally fine depending on what you're doing.)
When your architecture grows and you start needing to synchronize configuration across multiple applications, or have non-technical users change flags, then you've probably outgrown hardcoding and need to consider building (or buying) a service. You probably don't want product owners trying to set an env var like `FeatureManagement__NewFeature__EnabledFor__0__Name`. Azure App Configuration is the go to if you're willing to use Azure - it provides a nice UI for editing flags instead of dealing with JSON. Or, I've been building https://featureflags.app/ as an alternative provider if you prefer to steer clear of Azure. Its kinda the middle ground between hard coding and a full on platform like LaunchDarkly.
Simple, effective, cheap, easy to understand and manage. Not dependent on an external service, not dependent on a third-party.
The blogspam marketing behind them is so strong
Yep, right up there with Triplebyte back in the day and the drumbeat about "it's so hard to bill customers" and "JWT sucks"That covers a lot of uses of feature flags without the bloat.
> Yes to start out it’s 95% booleans. But then you want a rollout. And then you want some targeting rules. And then you want non booleans… maybe some json. Oo wouldn’t it be nice if the json could conform to a schema… and then eventually you are like damn I really want to change these without deploying. Or you want to read the same flag from multiple services.
First one got acquired. But also I wish I'd built it differently from the start (git based, other lessons learned), so started over.
It's true that it's a pretty simple thing, but you do want them to be of the utmost level of reliability. Its ideal if you can use feature flags & dynamic config (which I think of as the same problem) right at boot time. But then of course you have an external dependency to booting you app. So the reliability angle becomes imperative. So I do think the SaaS can play a good (or bad) role in overall reliability vs something you implement in house.
The other major advantage is telemetry. Once you have things toggling on and off, immediately engineers will get confused about whether something is on or off and why. Basic internal apps usually don't have any good telemetry / debugging. So a good SaaS should give you all kinds of ways to ensure that you can see what context is being passed around to the SDK to do their eval. Sanity checking etc.
My og experience was at HubSpot where we used the hell out of dynamic config. Once you get used to things like instantly targeting debug log levels for a single class to a single org its hard to go back.
There can be whole UIs and tooling and infrastructure to manage around them and that’s what the sass offer
You need to find every occurrence of flag4 and remove references to it. So that boolean expression would need to have `&& flag4` removed.
- one team uses feature flag for product gating. Feature flag service goes down. Users temporarily got locked out of the features they paid for.
- one team uses feature flag for dynamic pricing by leveraging targeting rules (how hard is it to write a bunch of if else in code?). It’s evaluated against all users, even if they are not active (for analysis reasons). Feature flag service charges by MAU. We have millions of users. Our feature flag service bill is now 6 digits per year.
- one team uses feature flag as literal json store instead of a proper db (god knows why). Someone updated the value but the “schema” is wrong. Shit breaks.
That could happen if they used a separate service. In this case you need to build defaults; open or closed, as the case demands.
> - one team uses feature flag for dynamic pricing by leveraging targeting rules (how hard is it to write a bunch of if else in code?). It’s evaluated against all users, even if they are not active (for analysis reasons). Feature flag service charges by MAU. We have millions of users. Our feature flag service bill is now 6 digits per year.
This is a valid reason if you don't own the service, as it seems you don't. If you did, you should have asked your internal customers what their requirements were.
> - one team uses feature flag as literal json store instead of a proper db (god knows why). Someone updated the value but the “schema” is wrong. Shit breaks.
Separating the services would not help here. They are simply using the wrong tool.
An example: at my current job, we’re using LaunchDarkly for config management, not just rollout. This means an incident on LaunchDarkly didn’t just stop a rollout, it caused the system behaviour to change for a large number of customers as it reverted to the default.