Cover image: US treasury 1920s.
A customer told me: “In GitOps, the number one failure mode is when people start treating Git like it is a database when it is a file system”. This is an extremely insightful comment because we need a reliable system of record for desired state, and those are normally databases.
So we must ask: should we also treat Git and its leading implementations as a database?
This customer runs large repos each with hundreds of users. For them the Git-database problem is so obvious that it doesn’t need spelling out. But if we don’t name and describe problems they never get solved. I am here to tell you that:
Treating Git as a Database leads to problems! How you represent, store and manage configuration is critical for software operations. Especially when you introduce generative AI and agents at any scale.
In this blog I’ll ask seven questions to help you see if you’re treating Git as a database, explain why each is a real business problem, and leave you with some tools to check it for yourself.
What Git likes and What Git doesn’t like
Let’s start with how we use Git today due to Config as Code.
We observe two elements:
-
In config-as-code, we represent configuration using source code and then we generate configuration data from that code
-
We store and manage it all as text in files in folders in Git code repos
We have covered the generation of configuration from templates in other posts. Let’s look at the storage element here: Git is a distributed versioned file store organised into folders like an operating system. The transaction model is optimized for teams working on different parts of one code base. Git is good at ordered, attributed, reviewable history of text. And we need ordering, attribution and review in a GitOps system of record.
The trouble starts when a repository becomes the operational store for a fleet. It then takes on the jobs of a database: keys, indexing, constraints, queries and transactions, potentially in bulk and at machine speed. Increasingly there may be machine data involved too. Operations may involve transactions across many data fields and objects at once. Git lacks the machinery to support this in a proper way that is fast, safe, scalable, reliable.
Without this machinery, modern teams have to create workarounds. This has happened before. Homebrew, CocoaPods, Go modules and Nixpkgs all started with a Git index and moved off it. They ‘hit a wall’ when too many users need to do something “ungitly” that Git doesn’t want to do.
Git has problems with managing configuration too. As well as the problem of how to store config, a complicating factor is the representation of config as code. Templating is fine, and so is Helm. What is missing is a store that can manage the mutation and fanout of one config code statement into multiple derived config data objects, without losing critical information. The rendered result may need to be addressed, queried or diffed. Today it is lost in CI and clusters.
OK now on to the fun part — the seven questions! I have divided them into three sections: Deployment issues, Promotions and Rollouts.
How do we map and approve deployments
1. Am I using filenames as primary keys?
Open your Git repos which store configuration. Do you have paths like this?
clusters/<org>-<env>-<region>.yaml
Databases have keys, with important uniqueness and referential properties. That filename is acting as a composite key. Your ‘config store in Git’ has directories instead of tables, and they have no type, no uniqueness check, no referential integrity. Ask: what happens when I rename a cluster? Or several clusters? Or a region? Or if someone adds a new app team?
A rename is now “a delete plus an insert”. The same fields are repeated inside every cluster and renaming one cluster touches about 20 files. This is because your ‘data’ is represented as text filenames.
Another issue is adding a value, and error handling:
-
In a database a new enum value failure is explicit and seen by consumers
-
In a folder scheme the lookup can miss, skip a layer, and report nothing
For example, suppose someone restructures org/<name>.yaml into org/<name>/<env>.yaml and some old lines stay in place. Suppose ignoreMissingValueFiles: true. Now some lookups resolve to nothing.
Clusters really do differ, and that is fine. In a database, we are equipped to handle variety and can describe variant clusters, apps, regions, teams, just as we represent bank accounts, cash, and customers in a business database.
Imagine describing a bank system as a set of files and folders, repos and orgs. You would have a sprawl of complex fields, text, documents and relations between them all. This breaks security and reliability. I don’t want to leave my money in a bank that doesn’t use a database, and I don’t think that modern (agentic) operations should be used without a database either.

2. Am I approving input queries, but deploying output values?
With agentic applications, every enterprise wants compliance and controls. That means approvals are cool again. Yay! But when you approve a PR, what exactly are you approving? Your review may happen before rendering. If so then what downstream values will get deployed, that depends on your review and approval of the upstream query?
For example: In traditional devops with config as code tools in Git, you may review a values change. What ships is a rendered manifest. But the render happens after review and before apply. That creates rendering and application risks that are not transparent or approved.
We have seen customers doing this across their whole estate. For example: over a thousand templated sources and no rendered manifests stored anywhere. Every approval was an “approval of an input”.
Argo CD has started on this with Source Hydrator, a clever workaround which is in beta. This renders manifests before committing them back to Git repos. That may fix what one Application renders to. It does not tell you how many Applications re-render when you change a shared file.
When is a promotion actually ‘done’
3. In my promotion flow, are transactional changes implemented as find-and-replace scripts?
A promotion is a sequence of approved changes to a system as explained here. Normally we would think of a promotion as a state change: this artifact, verified here, is now admitted there. In a database, that would be a sequence of ACID transactions.
In a file store, promotions involve updates to text files, via editing, and potentially some remapping of files into new folders. You might see a text edit repeated once per target group. Keeping a record that this happened, and can be reversed, or replicated, requires extra machinery. That is why CICD tools exist even when customers adopt GitOps successfully. Eg. Kargo already treats promotion as an object with an identity, stages and verification.
In one estate we looked at, a version string was repeated a median of seven times inside one definition, and over forty times in the worst. A partial replace is valid YAML, passes lint, and leaves the fleet on two versions with no record that anyone chose that.
Ask to see where it says this version was promoted to production, when, and who admitted it? Is it a script whose job is to ‘do the edit safely’? That’s a warning sign.
4. Can I consistently distinguish “in progress” changes from “abandoned”?
Do you know which of your charts are mid-promotion right now, and which ones stalled? Do you need to ask a colleague or check their text notes? This matters. A rollout that is 80% done and one we abandoned weeks ago can leave identical Git repositories. There is no start date, no target, no completion criterion and no expiry. Only version strings.
We have found that customers may be running two or more versions at once. An evil example: three versions spread across staging and production rollout waves. Tracing problems and their causes becomes very slow and painstaking at that point, unless you are willing to unleash AI on your live systems to get ‘a whole picture’. But why not fix the config management instead?
Tracing connections between files and rollouts
5. I edit one value. Do I use grep to calculate how many dependent values will change?
This is a classic blast radius problem which we have discussed in other blogs. In traditional devops we have Git files for config as code, templates and CI that then feeds into a set of clusters. The “Don’t Repeat Yourself” approach will mean that changes are handled in source and then rendered out through a tree of dependent manifests that each represent a real config.
What if one edited value produces as many changed values as our templating fans out to? Before you merge a change that may impact some or all of our whole fleet, can we say:
-
how many clusters does this touch
-
and how many values on each?
You see here the tension between one-to-many generation of config, from one DRY source to many targets, and the traditional GitOps model of one-one reconciliation between the live target system and the desired state ie. explicit or “Write Every Time” config.
In one example we saw: a one-line change to the global values file re-renders almost every Application in the fleet. The same change in one cluster’s file re-rendered a little over a hundred. The two diffs look the same, but one has a wildly larger blast radius.
Don’t forget: “density” multiplies the effect. Density is the number of impacted values per target. Blast radius is then: fan-out times density. This equals the number of targets a change reaches, times the number of values it changes on each.
In our example: The global file held about a hundred and twenty values. A single cluster’s file held more, nearly two hundred. Rewriting the global file could change over ten million live values. Rewriting a cluster’s file could change thirty thousand.
As we have seen, there is always a workaround. But is it safe to rely on patches, workarounds and so on when we are entering a new era of AI led automation and high scale?
In our example: The team built a blast-radius control: a hard CI check that caps a pull request at <N cluster files. The effect is to protect the per-cluster files; but these have the smallest reach and are owned by teams. The global file remains ungated without a natural owner. We are still trapped by “teams own paths in Git”, without blast control.
6. Do my controls reduce the rate of change, or make all changes safe?
In the Git world we set up all sorts of operational controls which help us work around its quirks. These can create a false sense of security.
List your controls and ask: Which of these makes a change safer, as opposed to rarer?
-
Files per pull request
-
Freeze windows
-
Patch cycles per year
-
Bot concurrency
-
Rollout waves
-
Limits on re-rendering
If every one of them limits how often you change things and none limits what one change does, you are rationing change. This is bad because cost per change still grows with fleet size. And, the change rate is rising because agents now write config. Total work is fleet size * change rate.
I.e: adding one more cluster costs its own work multiplied by every change made after it.
A system whose unit cost and unit risk both grow with scale has a ceiling: it is called your budget :-) And this is yet another reason why people use databases for business operations.
7. Does it always matter if a live resource has a different value from Git?
Last one! This is pretty subtle. GitOps says: make your live state reconcile with your desired state. But a large system could have many such loops reconciling at once. When is it ok to wait for convergence and when should you intervene? Using databases can give us more control.
Drift is the difference between the repo and the running system. That means a repo cannot show drift on its own; you need a live system to compare with.
Instead, your Git repo can show “what you allow”. In order to connect this up to live operations and management, once again we layer on workaround tools to capture decisions and explain them. For example: protecting a hand-edited production value; pruning resources from the cluster and desired state correctly; and of course letting automation loose: not just auto-scalers, but also AI. Put those rules in a database and link them to change management!
We have covered some of these topics on earlier blogs for example:
Does using AI and Machine-first Operations make this worse?
I think using AI to automate operations means that workarounds will break and create more complexity. To take a simple example: Git is not able to handle bulk ops at machine speed, and this isn’t fixable without really serious surgery. The world of “Git extensions” and “Post-Git stores” is upon us. ALL machines will speak Git in perpetuity, I suspect, but Git will not be the lingua franca of agentic operations, and “GitOps” will extend to new Git-like stores as expected.
Looking at the seven questions again: Each of these was survivable when people wrote config at human rates. Agents raise the rate of change, the reach of whoever holds the credentials, and the number of changes in flight at once. Human review can barely keep up. Example: A dependency bot supports unlimited generation, has a mandatory human merge, and a cap of three open pull requests. Unlimited machine writes, executed in groups of three at a time, but throttled if a person must check each one.
The seven questions again in short
-
Am I using filenames as primary keys?
-
Am I approving the query, but deploying the result?
-
Am I using find-and-replace as my transaction?
-
Am I using someone’s memory as my status column?
-
Am I using grep as my query engine?
-
Am I using rate limits as my constraints?
-
Am I using defaults as my drift policy?
Today, each of these is a database job being done by Git. Each one was a reasonable fix when people wrote config by hand, a few changes a day, and someone on the team remembered how it all fitted together. That is no longer how config gets written.
An agent needs three things a folder of files cannot give it. It needs to ask “which clusters run this version” and get an answer from a query, not a grep. It needs a missing reference or a missing value to be an error, not a silence. And it needs the reach of a change to be a number it can check before it acts, not a person who will look later.
What you can do this week
Turn off silent resolution. Set ignoreMissingValueFiles: false in CI only, render, and read the failures. You will find your dead layers in an afternoon. Then create the missing files or delete the dead lines, one chart at a time.
Put reach on the pull request. You probably already declare which paths affect which applications, for cache invalidation. Use the same data to count the targets each change touches, and post the number as a comment. Do not gate on it yet. Just stop your riskiest change looking the same as your safest one.
Keep the loop, swap the store, add the proof. Argo and Flux keep reconciling, and Helm keeps authoring. What changes is where the rendered result is held and what must be true before it moves. Once reach is a number, a quota can become a limit on what a change does, and cost per change follows what changed rather than the size of the fleet.
Find out where you stand
-
See the answers first. The example fleet is a small Argo CD repo with one planted instance of each of the seven, and an answer sheet. Two minutes.
-
Ask your own repo. Paste the seven-question prompt into Claude Code, Cursor, Codex or any agent with shell access to a clone of your config repo. It writes and runs its own scripts, reports only what they returned, and says NOT OBSERVED where the repo cannot answer. For Argo CD and Helm repos, scout runs the same seven questions as fixed checks in one read-only script, so you can repeat them exactly.
-
Talk to us. If the answers worry you, we would like to go through them with you (email)
Next time on this theme: a cost model so you can run the numbers yourself.
Keep Git for source code development, and give your fleet a GitOps-ready config database. Let us help make your software operations simpler, faster and safer — especially for agentic work.
–alexis

