[{"content":"\ntl;dr\ncoredns/dynapi v0.1.0 is out, published on 2026-10-01. It is an external CoreDNS plugin. Applications read, replace and delete A and AAAA record sets through a JSON HTTP API. The idea started as coredns/coredns#1822. I opened that PR on 2018-05-20. It was closed on 2018-09-29, when the coredns/dynapi repository was created. That repository stayed empty for eight years. The plugin in v0.1.0 is a new design, not a port of the old code. Why I needed it At SumUp we had what felt like an amazing idea: build our own take on ArgoCD. It was a Kubernetes-based deployment system for on-demand dev environments, and it ran on bare metal instances. It worked.\nEvery new environment needed DNS names, and it needed them fast. So we needed a DNS server that was reliable, simple to run and extensible. CoreDNS fit those three needs. One thing was missing: a way to register addresses dynamically. The obvious answer was an HTTP API, because any language and any tool can call one.\nThe 2018 story PR #1822 was titled \u0026ldquo;[WIP] Dynamic updates API with listen server\u0026rdquo;. It was an experimental REST API. A POST request with an A record added it to a zone managed by the file plugin.\nMy use case was narrow:\nCoreDNS is the main DNS server. I add records through an API. There is no database as a second source of truth. In-memory storage was good enough for me. We ran it in production from a little before 20 May 2018. A small daemon read DHCP leases and sent them to CoreDNS to register A records. It worked well, with one catch. Records created without a TTL got the default of 3600 seconds. That is far too long for short-lived environments, so the create and update requests got a TTL field.\nThe maintainers had a fair question. Miek Gieben asked why CoreDNS should provision records at all. Access control, users and permissions grow out of every API like this. The advice was to build it as an external plugin and put it at coredns/dynapi. John Belamaric suggested a different split: a Go \u0026ldquo;Writable\u0026rdquo; interface that plugins can implement, and separate plugins for the write protocols (DDNS, HTTPS, gRPC).\nThe same request was already open in #522 (\u0026ldquo;provide an API to manage host records\u0026rdquo;, 2017). In 2020 John Belamaric listed what a proper solution needs: a \u0026ldquo;writable\u0026rdquo; interface for backends, a plugin that exposes the REST API, and authentication that works across API plugins. #2073 asked for the same thing from the Ansible side, and #2517 was closed as unsupported, although a maintainer later said DNS UPDATE is \u0026ldquo;hard to do right, but I\u0026rsquo;m not against it\u0026rdquo;.\nThe repository was created in September 2018 and the PR was closed. Then my part stalled. COVID happened, I am bad at GitHub notifications, and I decided nobody needed the project. I never moved the code.\nI was wrong about the last part. For years, people kept asking in the closed PR. Some found the empty repository. Some asked why CoreDNS has no REST API. Some discussed which record types to support and whether Envoy\u0026rsquo;s xDS is a better model.\nOn 2026-10-01, out of nowhere, I checked again and saw the idea was still relevant. And here we are.\nWhat v0.1.0 does dynapi runs inside CoreDNS and opens its own HTTP listener. The default address is 127.0.0.1:8080, and only loopback addresses are allowed.\nThere is one URL shape:\n/v1/zones/{zone}/records/{name}/{type} The type is A or AAAA. Names must be full names inside the configured zone.\nGET returns the exact stored set, for example {\u0026quot;ttl\u0026quot;:60,\u0026quot;addresses\u0026quot;:[\u0026quot;192.0.2.10\u0026quot;]}. PUT replaces the whole set in one atomic step. DELETE removes the set and returns 204, also when the set does not exist. PUT is strict. It needs Content-Type: application/json, an explicit TTL from 0 to 2147483647, and 1 to 256 addresses of the right family. Duplicate addresses are removed. Bodies are limited to 64 KiB. Unknown fields and duplicate keys are rejected. The OpenAPI specification is generated from the Go models, and make verify checks that it is current.\ncurl -X PUT http://127.0.0.1:8080/v1/zones/example.org/records/host.example.org/A \\ -H \u0026#34;Authorization: Bearer $DYNAPI_TOKEN\u0026#34; \\ -H \u0026#39;Content-Type: application/json\u0026#39; \\ -d \u0026#39;{\u0026#34;ttl\u0026#34;:60,\u0026#34;addresses\u0026#34;:[\u0026#34;192.0.2.10\u0026#34;,\u0026#34;192.0.2.11\u0026#34;]}\u0026#39; dynupdate: the piece that showed up while I was away In 2018 CoreDNS had no way to change records at runtime, and I wrote my own. In 2026 it has one. The dynupdate plugin landed in CoreDNS itself in PR #8520. Contributor houyuwushang wrote it and Yong Tang merged it in September 2026. It speaks the standard protocol for this job, RFC 2136 DNS UPDATE, so nsupdate and DHCP servers such as Kea can already talk to it. The work started from #6254 (\u0026ldquo;Support DDNS\u0026rdquo;, open since 2023), and two core changes came first: #8469 lets the server accept UPDATE messages and #8471 exposes the validated TSIG identity to plugins.\nHow it works:\nOne writable zone per server block. You give it a seed zone file. The plugin never modifies that file. Persistence is optional. With database, every update goes to an embedded bbolt file before the new snapshot is visible and before the client sees success. Without it, updates live in memory and a restart loses them. The first access creates the database from the seed. After that, the database is the source of truth. Authentication is TSIG. The tsig plugin validates the signature. dynupdate never sees the secret. Authorization is explicit. Every change must match an allow KEY NAME TYPE rule. Nothing is allowed by default. The protocol is complete enough. It supports RFC 2136 prerequisites, add and delete operations, CNAME and apex SOA and NS rules, and automatic SOA serial updates. AXFR reads the current snapshot, and a successful change sends a best-effort NOTIFY. Caching is handled. The cache plugin bypasses dynamic zones, so authoritative answers are always current. Limits are stated. No DNSSEC records, no IXFR, no multi-primary replication. The README calls the plugin experimental and says it is meant for small zones that change rarely. That last point fits my own use case well. It also answers Miek\u0026rsquo;s old worry. CoreDNS has a place for write permissions and persistence now, and the HTTP layer can stay thin.\nHow it works dynapi has no record store of its own. It sends signed DNS requests to the same CoreDNS server over pooled loopback TCP connections.\nGET is a signed AXFR (zone transfer). PUT and DELETE are DNS UPDATE transactions. dynupdate owns the zone database. It handles persistence, atomic updates and write permissions. tsig authenticates the DNS requests. transfer allows the AXFR. There are two separate credentials. HTTP clients use a bearer token of at least 32 characters. It permits reads across the whole zone. The DNS side uses a TSIG key (HMAC-SHA256), and dynupdate allow rules limit what that key can write.\nIn the code, a PUT becomes one DNS UPDATE message:\nA prerequisite asserts that no CNAME exists at the name. RFC 2136 would silently ignore an add next to a CNAME, so this check keeps the HTTP response honest. The first update deletes the whole record set of that type. The next updates add the new addresses with the requested TTL. The server applies all of it as one transaction, so readers never see a half-replaced set. DELETE is the same message without the prerequisite and the adds. GET sends a signed AXFR and scans the transfer for the one name and type. The pool reuses TCP connections, and a connection that failed or was cancelled is discarded.\ndynupdate answers are mapped to HTTP status codes. A failed prerequisite is 409, a refused update is 403, and any other rejection is 502.\nThis is a complete Corefile:\nexample.org:1053 { bind 127.0.0.1 dynapi 127.0.0.1:8080 { token_env DYNAPI_TOKEN upstream 127.0.0.1:1053 identity update-key.example.org. secret_env DYNAPI_TSIG_SECRET } tsig { secret update-key.example.org. {$DYNAPI_TSIG_SECRET} require_opcode UPDATE require AXFR } transfer { to 127.0.0.1 } dynupdate { file example.org.zone database example.org.db allow update-key.example.org. host.example.org. A AAAA } } Why signed DNS and not a record store This design answers Miek\u0026rsquo;s 2018 concern. dynapi does not own users, permissions or storage. dynupdate and tsig already do that work, and dynapi only translates HTTP into signed DNS.\nADR 0001 records the decision and what it rejects:\nIts own record store. It would duplicate persistence, permissions and authoritative DNS behavior. Direct calls into dynupdate. That needs an upstream interface with defined transactions, permissions and lifecycle rules. Normal DNS queries for reads. They can expand wildcards and follow CNAMEs, so they cannot return the exact stored set. AXFR can. A web framework. Three HTTP operations do not need one, so dynapi uses net/http. PUT also checks for a conflicting CNAME in the same transaction.\nThe design has costs. GET scans the whole zone. Connections come from a pool limited by max_requests, which defaults to 32. Writes are never retried, because a change may commit before the response is lost.\nWhat is not done Only A and AAAA records work. The 2018 thread asked for more types, and Miek wanted storage that does not care about record types. That is future work. It needs Go 1.27 or newer. It also needs dynupdate and tsig from a pinned CoreDNS revision, so the development build is a CoreDNS binary built from the dynapi repository. Corefile reloads are rejected. Restart CoreDNS to change the configuration. Conditional writes (record revisions) are still open. Next: 0.2.0 The loopback DNS bridge works, but it is a detour. An HTTP request turns into a signed DNS message, goes over TCP to the same process, gets parsed and verified, and only then reaches the zone. A GET transfers the whole zone to read one record set.\nADR 0002 proposes the fix for 0.2.0. dynapi would call the zone provider directly. That removes the TCP hop, the full-zone reads and the TSIG key from dynapi. It is also the \u0026ldquo;Writable\u0026rdquo; interface John Belamaric described in 2018.\nI am not the first to ask. #7259 is an open proposal for a REST API plugin with a backend-agnostic api.Backend interface. Its author argues that the interface should live in the main CoreDNS project and that backends such as etcd should implement it. #7858 asks for a plugin that manages records through API or gRPC. Both point at the same gap.\nThe change in CoreDNS is needed first, because today plugins have no shared way to be written to:\nA Go interface to read, replace and delete stored record sets, with explicit errors. It starts with exact A and AAAA sets. Provider lookup by zone. CoreDNS would find the writable provider for a zone and manage its startup and shutdown. A provider that implements it. The interface alone stores nothing. dynupdate can be the first one. Clear rules for permissions, for when a write counts as durable, and for how cached answers are invalidated. Reload behavior. A Corefile reload must replace providers without sending requests to a retired zone. This also fixes the \u0026ldquo;reloads are rejected\u0026rdquo; limit above. Record revisions, so conditional writes can stop two clients from overwriting each other. The HTTP API stays the same during the move. The signed DNS adapter from v0.1.0 stays until the interface exists. Whether the writable provider ships with CoreDNS or as an optional plugin is still an open question. The first step is to agree on the interface upstream in the CoreDNS repository, before dynapi changes.\nTry it make builds CoreDNS with dynapi, dynupdate and tsig. Run ./coredns -plugins to check that all three are in the binary.\nTo add dynapi to another CoreDNS build, put this line in plugin.cfg after acl:\ndynapi:github.com/coredns/dynapi/plugins/dynapi Then run:\ngo get github.com/coredns/dynapi/plugins/dynapi@v0.1.0 go generate coredns.go go build -o coredns . The examples/ directory has a starter Corefile and an annotated Go client.\nSorry to all the patient people waiting. :D Now we\u0026rsquo;ll follow normal OSS cadence. Any feedback, issues and PRs are welcome in the GitHub repo.\n","permalink":"https://syndbg.dev/posts/2026-10-02-coredns-dynapi-v0-1-0/","summary":"\u003cp\u003e\u003cimg alt=\"dynapi demo\" loading=\"lazy\" src=\"https://github.com/coredns/dynapi/raw/v0.1.0/docs/preview.gif\"\u003e\u003c/p\u003e\n\u003cp\u003etl;dr\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"https://github.com/coredns/dynapi/releases/tag/v0.1.0\"\u003ecoredns/dynapi v0.1.0\u003c/a\u003e is out, published on 2026-10-01.\u003c/li\u003e\n\u003cli\u003eIt is an external CoreDNS plugin. Applications read, replace and delete A and AAAA record sets through a JSON HTTP API.\u003c/li\u003e\n\u003cli\u003eThe idea started as \u003ca href=\"https://github.com/coredns/coredns/pull/1822\"\u003ecoredns/coredns#1822\u003c/a\u003e. I opened that PR on 2018-05-20. It was closed on 2018-09-29, when the \u003ccode\u003ecoredns/dynapi\u003c/code\u003e repository was created.\u003c/li\u003e\n\u003cli\u003eThat repository stayed empty for eight years. The plugin in v0.1.0 is a new design, not a port of the old code.\u003c/li\u003e\n\u003c/ul\u003e\n\u003chr\u003e\n\u003ch2 id=\"why-i-needed-it\"\u003eWhy I needed it\u003c/h2\u003e\n\u003cp\u003eAt SumUp we had what felt like an amazing idea: build our own take on \u003ca href=\"https://argo-cd.readthedocs.io/\"\u003eArgoCD\u003c/a\u003e. It was a Kubernetes-based deployment system for on-demand dev environments, and it ran on bare metal instances. It worked.\u003c/p\u003e","title":"coredns/dynapi v0.1.0: an HTTP API for CoreDNS records, 8 years later"},{"content":" tl;dr\nNotes: 0.2.0 at github.com/syndbg/onetui/releases. Link back to the 0.1.0 post. What\u0026rsquo;s new, at a glance Notes: Every data source can now write, from the same query editor: PostgreSQL, Kafka, NATS, DynamoDB, RabbitMQ, Qdrant. New data source: ScyllaDB and Cassandra, through one cql connection kind. Syntax highlighting in the editor, per data source language. Multi-line editor (Shift+Enter), query history (Ctrl-P / Ctrl-N, Shift+H), optional persistent history. Confirmation before a query runs, on by default (ask_for_query_confirm). Connection errors in a popup. Inline username= / password= credentials in connection config (#41). The 0.1.0 post named four sources. 0.1.0 already shipped DynamoDB and RabbitMQ as read-only too. 0.2.0 makes seven. PostgreSQL: SQL with highlighting and a confirmation step ScyllaDB and Cassandra: one CQL provider Connection errors in a popup Kafka: PRODUCE and CONSUME NATS JetStream: the same verbs RabbitMQ: verbs instead of URL-encoded paths DynamoDB: PartiQL through ExecuteStatement Qdrant: HTTP requests What 0.2.0 is about Getting 0.2.0 out was mostly time spent thinking. Not typing, thinking. How do you make the developer experience feel the same when one data source speaks SQL, another speaks HTTP and a third is a log of bytes with no query language at all?\nFrom read-only to writes In the 0.1.0 post I said read-only comes first, because writes raise the stakes: wrong target, wrong environment, data loss. I also said autocomplete would come before writes. Well, writes shipped first. Autocomplete didn\u0026rsquo;t make it.\nWhy I\u0026rsquo;m fine with that: autocomplete helps you type the right thing. It doesn\u0026rsquo;t stop you from running the wrong thing against the wrong cluster. What does help:\nBefore anything runs, a popup tells you which connection and which target you\u0026rsquo;re about to hit. It\u0026rsquo;s on by default. If you find it annoying, ask_for_query_confirm = false turns it off. Your call. Every write tells you what happened: applied, rejected or unknown. Nothing retries on its own. Ever. If OneTUI says \u0026ldquo;unknown\u0026rdquo;, go look before you hit Enter again. A resend can duplicate a Kafka record, bump a Cassandra counter twice or append to a list twice, and only you know if that\u0026rsquo;s fine. A rejected query doesn\u0026rsquo;t kill your connection. You fix the typo and carry on. And the boring one that still matters most: use credentials that can\u0026rsquo;t do more than they should.\nOne editor, many languages Every data source gets the same editor and the same flow. Press e, type, confirm, run, look at the result. Ctrl-P brings back what you ran before, Ctrl-C cancels. Same keys whether it\u0026rsquo;s Postgres or RabbitMQ.\nThe TUI doesn\u0026rsquo;t try to guess if your query reads or writes. It hands the text to the data source\u0026rsquo;s provider, and the provider decides. OneTUI never has to understand SQL, CQL and PartiQL to know what you meant, and that\u0026rsquo;s the reason it scales.\nThe rule I landed on is simple:\nIf the data source already has a query language, you write that language, as is. SQL for Postgres, CQL for ScyllaDB and Cassandra, PartiQL for DynamoDB. If it doesn\u0026rsquo;t, you write a verb line that names the target, then a blank line, then a body. That\u0026rsquo;s it. That\u0026rsquo;s the whole DSL.\nData source What you type Example PostgreSQL SQL SELECT * FROM demo.customers ScyllaDB, Cassandra CQL SELECT * FROM ks.events WHERE bucket = 0 DynamoDB PartiQL, inside a JSON operation {\u0026quot;operation\u0026quot;: \u0026quot;ExecuteStatement\u0026quot;, ...} Qdrant HTTP POST /collections/demo/points/scroll Kafka Verbs CONSUME orders/0 offsets 123..250, PRODUCE orders NATS Verbs CONSUME EVENTS orders.* seq 1..500, PRODUCE EVENTS orders.created RabbitMQ Verbs PUBLISH / amq.default demo, DECLARE queue / demo, RAW GET /api/overview I didn\u0026rsquo;t start there. Kafka and NATS first took a JSON object, and the keys in it decided what happened. That was fine for one operation and fell apart at two. It also had nowhere to say which topic or stream you meant. So the query only made sense from the screen you typed it on.\nRabbitMQ was the other lesson. I first passed GET /api/... straight through to the management API, same as Qdrant. Then you meet the default vhost. It\u0026rsquo;s literally named /, so in a URL it has to be %2F. Every path turned into /api/queues/%2F/demo, and if you got the encoding wrong, RabbitMQ gave you a bare 404 and a shrug. Now you type GET / demo and OneTUI does the encoding. For the long tail of endpoints the verbs don\u0026rsquo;t cover, RAW is still there.\nQdrant keeps plain HTTP, since its REST API is already what people know.\nA verb line puts the target up front, so a query means the same thing wherever you run it. And since OneTUI knows which view you came from, it can prefill that line for you.\nI think this is the approach that scales best across data sources while staying relaxed about it. There\u0026rsquo;s no grand universal query language to learn, and no pretending a Kafka topic is a SQL table. I\u0026rsquo;m fairly sure even an RPC-only database like TigerBeetle would fit. It has no query language, just a handful of operations, which is exactly what verbs are for:\nLOOKUP accounts 1 2 3 CREATE transfers {\u0026#34;id\u0026#34;: 1, \u0026#34;debit_account_id\u0026#34;: 1, \u0026#34;credit_account_id\u0026#34;: 2, \u0026#34;amount\u0026#34;: 10, \u0026#34;ledger\u0026#34;: 1, \u0026#34;code\u0026#34;: 1} Not built, just a sketch. But it slots in without changing anything else.\nSyntax highlighting, and why not tree-sitter Once the language question was settled, highlighting was almost free. Every data source uses one of three syntaxes: SQL (plus its own extra keywords), JSON, or a verb line with a body.\nMy first instinct was tree-sitter, because that\u0026rsquo;s what everyone uses. I tried it, and a few other options, on the same four half-typed queries. tree-sitter\u0026rsquo;s SQL grammar colored 0 as a string and tags CONTAINS 'a' as one big string. The PartiQL parser keeps its lexer private and drags in 56 crates. So I went with the sqlparser tokenizer for SQL and CQL, the jsonc-parser scanner for JSON, and a tiny lexer of my own for the verb line.\nWhat I cared about:\nHalf-typed input keeps its color. An unfinished string stays a string instead of the whole line going gray. Message payloads don\u0026rsquo;t get colored. They\u0026rsquo;re sent as raw bytes, so pretending they\u0026rsquo;re JSON would lie to you. It reuses each theme\u0026rsquo;s existing colors, so all ten themes just work. What\u0026rsquo;s next in 0.3.0 Autocomplete and my take on self-documenting CLI that may also have an MCP and skills. Custom keybinds? More themes?\nCode is at github.com/syndbg/onetui.\n","permalink":"https://syndbg.dev/posts/2026-09-28-onetui-0-2-0/","summary":"\u003c!--\nResearch draft. Headings, GIFs and facts are in place. Bullets are notes for the prose, not final text.\nSources: onetui v0.1.0..origin/main, ADR-0012 to ADR-0019, README \"Features and datasource support\".\n--\u003e\n\u003cp\u003e\u003cimg alt=\"OneTUI logo\" loading=\"lazy\" src=\"https://github.com/syndbg/onetui/raw/main/docs/assets/onetui-logo.png\"\u003e\u003c/p\u003e\n\u003cp\u003etl;dr\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eNotes: 0.2.0 at \u003ca href=\"https://github.com/syndbg/onetui/releases\"\u003egithub.com/syndbg/onetui/releases\u003c/a\u003e. Link back to the \u003ca href=\"/posts/2026-09-20-onetui-one-terminal-every-data-source/\"\u003e0.1.0 post\u003c/a\u003e.\u003c/li\u003e\n\u003c/ul\u003e\n\u003chr\u003e\n\u003ch2 id=\"whats-new-at-a-glance\"\u003eWhat\u0026rsquo;s new, at a glance\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003eNotes:\n\u003cul\u003e\n\u003cli\u003eEvery data source can now write, from the same query editor: PostgreSQL, Kafka, NATS, DynamoDB, RabbitMQ, Qdrant.\u003c/li\u003e\n\u003cli\u003eNew data source: ScyllaDB and Cassandra, through one \u003ccode\u003ecql\u003c/code\u003e connection kind.\u003c/li\u003e\n\u003cli\u003eSyntax highlighting in the editor, per data source language.\u003c/li\u003e\n\u003cli\u003eMulti-line editor (Shift+Enter), query history (Ctrl-P / Ctrl-N, Shift+H), optional persistent history.\u003c/li\u003e\n\u003cli\u003eConfirmation before a query runs, on by default (\u003ccode\u003eask_for_query_confirm\u003c/code\u003e).\u003c/li\u003e\n\u003cli\u003eConnection errors in a popup.\u003c/li\u003e\n\u003cli\u003eInline \u003ccode\u003eusername=\u003c/code\u003e / \u003ccode\u003epassword=\u003c/code\u003e credentials in connection config (#41).\u003c/li\u003e\n\u003cli\u003eThe 0.1.0 post named four sources. 0.1.0 already shipped DynamoDB and RabbitMQ as read-only too. 0.2.0 makes seven.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3 id=\"postgresql-sql-with-highlighting-and-a-confirmation-step\"\u003ePostgreSQL: SQL with highlighting and a confirmation step\u003c/h3\u003e\n\u003cp\u003e\u003cimg alt=\"PostgreSQL query with syntax highlighting and confirmation\" loading=\"lazy\" src=\"/images/onetui-0.2.0/postgres.gif\"\u003e\u003c/p\u003e","title":"OneTUI 0.2.0: writes, CQL and one way to query everything"},{"content":"\ntl;dr Code is at github.com/syndbg/onetui. If the idea of one consistent TUI across your databases and queues sounds useful to you too, I\u0026rsquo;d like to hear about it.\nOver the years I\u0026rsquo;ve used DBeaver, DataGrip, psql, and whatever Kafka CLI happened to be the flavor of the month. Each one does its job, more or less. None of them feel like the same tool. Keybinds differ. Panels differ. Some support the data source I need that day, some don\u0026rsquo;t. Some barely work.\nWhat\u0026rsquo;s missing across all of them isn\u0026rsquo;t features or polish. It\u0026rsquo;s one interaction model I can get used to and use consistently, whether I\u0026rsquo;m looking at Postgres rows or a Kafka topic. And why not a few more data sources too?\nStanding on prior tools OneTUI isn\u0026rsquo;t a from-scratch idea and I don\u0026rsquo;t want to pretend it is. It borrows on purpose: K9s\u0026rsquo;s resource and navigation model and adds a thin TUI layer over the raw protocol. DBeaver and DataGrip\u0026rsquo;s breadth of data source support (minus the inconsistency), the Kafka/Redpanda support and expectations for various schema sources and decoding - Buf, Proto, Avro, Schema Registry. The point isn\u0026rsquo;t novelty. It\u0026rsquo;s picking the parts that already proved themselves and dropping the parts that didn\u0026rsquo;t, inside one consistent shell.\nMost useful tools work this way. Few are invented whole.\nK9s, and why it\u0026rsquo;s the reference point I admire K9s for lasting this long as a consistent tool. It has rough edges (looking at you slow timeouts when starting a new session), one set of keybinds, one mental model for every Kubernetes resource. People don\u0026rsquo;t just tolerate K9s, they reach for it first. That\u0026rsquo;s where I wish to get.\nThe idea: OneTUI My goal is not to sound like the XKCD comic about standards and develop one more tool to solve all problems before the next one comes and tries to do the same. :D\nWith the bold claim \u0026ldquo;OneTUI is one terminal UI for every database, message queue, or data source you connect to. The goal is consistent resources, consistent keybinds, and a consistent interaction model, no matter which data source is on the other end.\u0026rdquo;\nToday that\u0026rsquo;s Postgres, Kafka, NATS, and Qdrant, each with a working provider:\nPostgres: browse tables, inspect a row field by field, run a native query above the results. Kafka: browse topics, follow live records, decode Protobuf and Avro payloads against a schema registry automatically. NATS: browse subjects and follow live messages, same resource model as Kafka. Qdrant: browse collections and points, and inspect cluster consensus state through the same resource browser used for a Postgres table. None of these get a separate UI, all are integrated in the same consistent TUI. They all sit behind the same keybinds, the same panel layout, the same way of drilling from a resource list into a single item. That\u0026rsquo;s the entire goal: the tool underneath can be anything, as long as what you see and press stays the same.\nRight now, that consistency is easiest to hold onto in read-only mode. So that\u0026rsquo;s where OneTUI starts.\nWhy read-only first Write support raises the stakes immediately: wrong target, wrong environment, data loss. None of that is worth risking before the UI and the connection model have proven themselves. Read-only keeps the scope small on purpose. It\u0026rsquo;s a starting constraint, not a missing feature. Once the core model is solid, write support has a foundation to build on instead of a rushed one to patch.\nOne of my immediate next goals is, unless there\u0026rsquo;s my personal interest of supporting ScyllaDB/Cassandra first, to add auto-complete for the editor and that\u0026rsquo;ll be a prerequisite for writing.\nTech and architecture OneTUI is written in Rust, using ratatui for the terminal UI.\nThe Rust pick has two reasons. Personally, I want to learn something new. Been writing Go for 12 years, so it\u0026rsquo;s enough for me to look at a different language for a while. No need for me to tell you why Rust over Go, etc.\nRust, Ratatui, Tokio, plus all the necessary dependencies to make a data source connection work.\nWhile I am quite proficient with Go, I am writing production Rust code outside of OneTUI too. I am far from calling myself proficient at the same level and that\u0026rsquo;s why I use LLMs to help write the Rust.\nLanguage-aside, the code is always human reviewed. Guardrails this, guard rails that, it\u0026rsquo;s human reviewed. No skipping that part. I carry my expertise using all these databases and message queues too.\nAdding a data source Every data source implements the same two traits: Provider describes it (connection fields, which resources it exposes, whether it supports live following or native queries) and builds a configured Executor; Executor does the work (check, fetch_page, status, and capability-gated follow_page / query_page for sources that support them). The TUI never talks to Postgres or Kafka directly. It talks to Executor, always the same way.\nThis isn\u0026rsquo;t a dynamic plugin system, by design. OneTUI uses static enum dispatch: BuiltinProvider and BuiltinExecutor enums, one variant per data source, matched exhaustively. No Box\u0026lt;dyn Provider\u0026gt;, no async-trait, no boxed futures at that seam, just native impl Future returns the compiler can check and inline. The tradeoff is explicit: a new source needs a new crate, a new enum variant, a forwarding arm in the match, and a catalog entry, then a rebuild. What you get back is a compile error the moment dispatch is incomplete, instead of a runtime surprise from a source nobody finished wiring up.\nflowchart LR subgraph Sources[\"Data sources\"] PG[(Postgres)] KF[(Kafka)] NT[(NATS)] QD[(Qdrant)] NEW[(\"New source\")] end subgraph Crates[\"Provider crates\"] PGC[onetui-postgres] KFC[onetui-kafka] NTC[onetui-nats] QDC[onetui-qdrant] NEWC[\"onetui-\u0026lt;new\u0026gt;\"] end PG --- PGC KF --- KFC NT --- NTC QD --- QDC NEW -.implements.-\u003e NEWC PGC --\u003e BP[\"BuiltinProvider Postgres(...) | Kafka(...) | Nats(...) | Qdrant(...) | ...\"] KFC --\u003e BP NTC --\u003e BP QDC --\u003e BP NEWC -.-\u003e|\"new enum variant + forwarding arm\"| BP BP -- \"configure()\" --\u003e BE[\"BuiltinExecutor same variants, exhaustive match per method\"] BE --\u003e TUI[\"TUI (ratatui) same resource browser, same keybinds\"] flowchart LR subgraph Sources[\"Data sources\"] PG[(Postgres)] KF[(Kafka)] NT[(NATS)] QD[(Qdrant)] NEW[(\"New source\")] end subgraph Crates[\"Provider crates\"] PGC[onetui-postgres] KFC[onetui-kafka] NTC[onetui-nats] QDC[onetui-qdrant] NEWC[\"onetui-\u0026lt;new\u0026gt;\"] end PG --- PGC KF --- KFC NT --- NTC QD --- QDC NEW -.implements.-\u003e NEWC PGC --\u003e BP[\"BuiltinProvider Postgres(...) | Kafka(...) | Nats(...) | Qdrant(...) | ...\"] KFC --\u003e BP NTC --\u003e BP QDC --\u003e BP NEWC -.-\u003e|\"new enum variant + forwarding arm\"| BP BP -- \"configure()\" --\u003e BE[\"BuiltinExecutor same variants, exhaustive match per method\"] BE --\u003e TUI[\"TUI (ratatui) same resource browser, same keybinds\"] ⤢ Adding a source means writing a new crate against Provider/Executor, then wiring one enum variant into the catalog: no changes to the TUI, no changes to the other providers\u0026rsquo; code. That\u0026rsquo;s the part of the architecture the whole premise leans on.\nHow the TUI is put together The render loop never blocks on network calls. Each connection gets a Worker running its Executor on its own async task; the TUI thread only exchanges messages with it.\nflowchart TB App[\"App owns View state, dispatches Requests\"] Worker[\"Worker (per connection) tokio task wrapping one Executor\"] Exec[\"Executor talks to Postgres / Kafka / NATS / Qdrant / ...\"] UI[\"ui.rs renders View with ratatui\"] App -- \"Request (mpsc)\" --\u003e Worker Worker -- \"spawns\" --\u003e Exec Worker -- \"Page / Status (mpsc + watch)\" --\u003e App App --\u003e UI UI -- \"keypress\" --\u003e App flowchart TB App[\"App owns View state, dispatches Requests\"] Worker[\"Worker (per connection) tokio task wrapping one Executor\"] Exec[\"Executor talks to Postgres / Kafka / NATS / Qdrant / ...\"] UI[\"ui.rs renders View with ratatui\"] App -- \"Request (mpsc)\" --\u003e Worker Worker -- \"spawns\" --\u003e Exec Worker -- \"Page / Status (mpsc + watch)\" --\u003e App App --\u003e UI UI -- \"keypress\" --\u003e App ⤢ A Request goes out, a Page or a ConnectionStatus comes back on a channel, the view updates, ratatui redraws. Same loop regardless of which provider is on the other end, which is what keeps the keybinds and panel behavior identical across every data source.\nA TUI still needs personality Consistency doesn\u0026rsquo;t mean the tool has to look the same for everyone. OneTUI ships ten built-in color themes you can switch between, so the shell that\u0026rsquo;s identical across every data source doesn\u0026rsquo;t have to be visually dull across every terminal.\nPlus Solarized, Nord, Dracula, Tokyo Night, One Dark, Rosé Pine, Monokai, and Flexoki.\nCustom keybinds aren\u0026rsquo;t there yet. The current set is fixed. But the architecture makes adding remapping straightforward, there are a few reasonable ways to do it, and I haven\u0026rsquo;t had to pick one yet because nobody\u0026rsquo;s asked. If enough people want it, it gets prioritized.\nWhere this gets fun to use: Kafka messages encoded in Protobuf or Avro, decoded live against a schema registry, with the schema auto-detected. No manual schema picking, no separate deserializer bolted on the side.\nAuto-detection works well. Rendering still has minor rough edges I haven\u0026rsquo;t polished out yet, worth saying plainly rather than glossing over. But this is the part that turns the \u0026ldquo;consistent TUI across data sources\u0026rdquo; pitch into something I can point at and say: it already works.\nQdrant\u0026rsquo;s consensus state, in the same resource browser used for a Postgres table:\nWhere it stands, and where it\u0026rsquo;s going OneTUI is publicly released as I write this blog post. This is not a big deal, it\u0026rsquo;s just another tool, I only want to make something useful.\nI\u0026rsquo;ve been sharing it with a few colleagues and friends. The feedback is positive. Any feedback, GitHub issue or time spent using it is greatly appreciated.\nNow where it is going? The architecture leaves room to grow: new data sources, new resource types. In practice that means almost any data source is fair game, whatever exposes rows, records, or resources through some protocol fits the same provider shape. As I mentioned earlier ScyllaDB/Cassandra makes sense for the sake of supporting the most popular protocols.\nStill. the shared goal across all of it isn\u0026rsquo;t the count of features. It\u0026rsquo;s a TUI people actually want to reach for, the same reason K9s wins over raw kubectl for a lot of people day to day: not more power, just less friction to use the functionality that\u0026rsquo;s already there.\nWill OneTUI replace psql and active tools used to query and modify data sourecs? Not today. psql has decades of trust and a query editor good developers already know from muscle memory. But that gap is closable: real autocomplete against the live schema, a proper multi-line editor instead of a single input line, history that\u0026rsquo;s searchable instead of scrollback. None of that requires reinventing the protocol underneath, it\u0026rsquo;s interface work. I think a TUI can beat a REPL at its own game if it takes the editing experience as seriously as the data browsing.\nCode is at github.com/syndbg/onetui. If the idea of one consistent TUI across your databases and queues sounds useful to you too, I\u0026rsquo;d like to hear about it.\n","permalink":"https://syndbg.dev/posts/2026-09-20-onetui-one-terminal-every-data-source/","summary":"\u003cp\u003e\u003cimg alt=\"OneTUI logo\" loading=\"lazy\" src=\"https://github.com/syndbg/onetui/raw/main/docs/assets/onetui-logo.png\"\u003e\u003c/p\u003e\n\u003cp\u003etl;dr Code is at \u003ca href=\"https://github.com/syndbg/onetui\"\u003egithub.com/syndbg/onetui\u003c/a\u003e. If the idea of one consistent TUI across your databases and queues sounds useful to you too, I\u0026rsquo;d like to hear about it.\u003c/p\u003e\n\u003chr\u003e\n\u003cp\u003eOver the years I\u0026rsquo;ve used DBeaver, DataGrip, psql, and whatever Kafka CLI happened to be the flavor of the month. Each one does its job, more or less. None of them feel like the same tool. Keybinds differ. Panels differ. Some support the data source I need that day, some don\u0026rsquo;t. Some barely work.\u003c/p\u003e","title":"OneTUI: one terminal, every data source"},{"content":"I\u0026rsquo;ve used z for over a decade. Not zoxide, not jump, not fasd, the original rupa/z, a few hundred lines of POSIX shell that tracks the directories you cd into and lets you jump back with a fuzzy pattern instead of a full path.\nAt some point I tried the newer alternatives. Each one changed something I didn\u0026rsquo;t ask it to change: a different ranking algorithm, a different database format, a new set of flags to relearn. None of them were bad tools. They just weren\u0026rsquo;t z. I wanted the exact behavior I already had muscle memory for, just without a shell interpreter parsing a database file on every invocation.\nSo I wrote zjyo a while back, a 1:1 Rust port. The only thing that changed is what\u0026rsquo;s running under the hood.\nRewriting working software is proof the original design was already good rupa/z is just good. The matching, the aging, the flags, none of it needed fixing, so I didn\u0026rsquo;t touch any of it. I just moved the same design into Rust and kept my hands off the parts that were already right. rupa wrote a man page instead of a friendly README, kept the whole thing to one shell script, and never bloated it with features nobody asked for. That\u0026rsquo;s taste. It\u0026rsquo;s rarer than it should be.\nMost of the OSS I run into these days wants to be everything. Plugins, config files, a dashboard, a roadmap. rupa/z never went there, and it\u0026rsquo;s still exactly as useful seventeen years later as it was on day one. We need more of that: people who build one small thing, get the shape of it right, and then leave it alone.\nWhy bother A rewrite for its own sake is a waste of time. At some point it was a matter of writing Rust for the fun and seeing how much handholding I need with an LLM to assist me. Turned out well, since I use it for a few years now.\nThe end result zjyo is under 800 lines of Rust across five files. database.rs handles persistence and matching. entry.rs is a plain struct with a frecency calculation. cli.rs wires up clap and dispatches. That\u0026rsquo;s the whole surface area, and most of what maintaining a small tool actually looks like: not new features, but noticing the gap between what you assumed was true and what was actually happening.\nThe bug that mattered most The tmux bug I hit while using tmux-continuum is worth mentioning. First, using tmux and tmux-continuum is something that I don\u0026rsquo;t plan to change in the next 10 years too. It just works. The issue is when you make an assumption that cd is going to always work and wondering hey why does zjyo do it differently, some things start to make sense, because it\u0026rsquo;s a class of bug that\u0026rsquo;s easy to dismiss as \u0026ldquo;user error\u0026rdquo; and hard to find by reading code. The database was updating, for every shell session I\u0026rsquo;d opened since editing .zshrc. But tmux-continuum restores sessions across machine restarts, and panes that survive a restore don\u0026rsquo;t re-source your rc files. Config and the cd wrapper are two different things, and only one of them updates when you edit a file.\nThe fix, using precmd_functions instead of overriding cd, happens to also route around the entire class of bug, since it\u0026rsquo;s additive rather than a redefinition. It doesn\u0026rsquo;t fix stale shells. Nothing can fix a process that\u0026rsquo;s already running with old code loaded. But it means the next time something like this happens, at least new panes won\u0026rsquo;t diverge from what you think your shell does without you noticing.\nWalking through it The original integration looked like this:\ncd() { builtin cd \u0026#34;$@\u0026#34; \u0026amp;\u0026amp; zjyo --add } Redefining cd as a shell function is the obvious approach, anyone who has skimmed a z integration script has seen this pattern. It works, right up until something else in your shell startup also defines cd. Whichever definition loads last wins, with no error and no warning. The earlier one just stops existing.\nThat wasn\u0026rsquo;t actually my bug though, no other plugin was fighting for cd in my case. The real problem was simpler and harder to see: some of my zjyo-tracked directories just weren\u0026rsquo;t showing up in z -l, despite me having cd\u0026rsquo;d into them dozens of times that day. The database file (~/.z) was clearly being written to. stat ~/.z showed a recent mtime, other directories were tracked fine, and zjyo --add worked perfectly when I ran it by hand in the same pane. So the binary was fine, the wrapper function was fine when invoked, and yet specific panes weren\u0026rsquo;t invoking it.\nThe panes in question all had one thing in common: they were tmux panes that survived a tmux-continuum restore, meaning they were spawned before I\u0026rsquo;d last edited .zshrc to add the zjyo integration. A shell process reads its rc file once, at startup. Editing ~/.zshrc after the fact does nothing to a shell that\u0026rsquo;s already running. The function definitions it loaded at spawn time are the only ones it has. tmux-continuum is built to survive machine restarts by serializing pane state and restoring it later, so a pane you\u0026rsquo;re typing into today can be running a shell process that\u0026rsquo;s been alive, uninterrupted, since before a config change you made weeks ago.\nConfirming this took nothing more than:\n# in a suspect pane type cd # cd is a shell builtin versus a freshly spawned pane:\ntype cd # cd is a shell function from /Users/syndbg/.zshrc If cd reports as a builtin instead of a function, that pane never picked up the wrapper. No amount of retrying z --add in that shell fixes it, because the function it would need to call doesn\u0026rsquo;t exist there.\nprecmd_functions gets around this. Every zsh prompt draw runs every function registered in that array:\n_zjyo_precmd() { (zjyo --add \u0026amp;) } precmd_functions+=(_zjyo_precmd) += means appending, not overwriting, so it can\u0026rsquo;t replace someone else\u0026rsquo;s hook the way redefining cd can. And because it hooks the prompt rather than the cd builtin specifically, it also picks up pushd, popd, and any script that changes directory without going through the user\u0026rsquo;s interactive cd at all. It doesn\u0026rsquo;t retroactively fix the stale panes sitting there with the old function loaded, those still need a source ~/.zshrc or a fresh pane. But new panes stop quietly diverging from what your config says.\nWhy keep doing this There\u0026rsquo;s no ecosystem reason to prefer zjyo over zoxide. zoxide has more users, more contributors. The reason to maintain your own small thing isn\u0026rsquo;t that it\u0026rsquo;s objectively better. It\u0026rsquo;s that you understand every line of it, you can fix what actually bothers you instead of filing an issue and waiting, and the maintenance itself is quite lean when there\u0026rsquo;s not much functionality and need for it, to begin with.\n800 lines and a decade-old database format don\u0026rsquo;t need a roadmap. They need someone willing to keep them that small.\n","permalink":"https://syndbg.dev/posts/2026-09-19-on-writing-small-useful-things-like-zjyo/","summary":"\u003cp\u003eI\u0026rsquo;ve used \u003ccode\u003ez\u003c/code\u003e for over a decade. Not \u003ccode\u003ezoxide\u003c/code\u003e, not \u003ccode\u003ejump\u003c/code\u003e, not \u003ccode\u003efasd\u003c/code\u003e, the original \u003ca href=\"https://github.com/rupa/z\"\u003erupa/z\u003c/a\u003e, a few hundred lines of POSIX shell that tracks the directories you \u003ccode\u003ecd\u003c/code\u003e into and lets you jump back with a fuzzy pattern instead of a full path.\u003c/p\u003e\n\u003cp\u003eAt some point I tried the newer alternatives. Each one changed something I didn\u0026rsquo;t ask it to change: a different ranking algorithm, a different database format, a new set of flags to relearn. None of them were bad tools. They just weren\u0026rsquo;t \u003ccode\u003ez\u003c/code\u003e. I wanted the exact behavior I already had muscle memory for, just without a shell interpreter parsing a database file on every invocation.\u003c/p\u003e","title":"On writing small useful things like zjyo"},{"content":"A Cloud Controller Manager is a bridge. On one side: a Kubernetes cluster that thinks in Nodes, Services, and Routes. On the other: whatever actually runs your machines and network - a hyperscaler like AWS, GCP, or Azure, or just as often a private cloud, a colocated rack, or a home lab. The CCM\u0026rsquo;s whole job is translation - when a Node joins, ask the infrastructure what it knows about that machine and translate it onto the Node object; when a Service wants a load balancer, go create one and wire the ingress IP back into status. Without it, kubectl get nodes never gets a region label and type: LoadBalancer never gets an IP - Kubernetes has no idea any of that infrastructure exists.\nI\u0026rsquo;ve spent a good chunk of time in this code, both building on top of kubernetes/cloud-provider directly and reading through what Hetzner\u0026rsquo;s and DigitalOcean\u0026rsquo;s CCMs do differently. This post covers the interface, the four control loops that drive it, the full node/route/service lifecycle, and the patterns worth stealing once you\u0026rsquo;ve got the basics working.\nWhat is a CCM? Cloud providers used to be compiled directly into kube-controller-manager - the same coupling problem CSI solved for storage. If you\u0026rsquo;ve ever run a non-cloud cluster and watched the logs fill up with a cloud provider that isn\u0026rsquo;t even yours failing to authenticate, that\u0026rsquo;s the symptom. The Kubernetes docs on cloud controller manager architecture cover the history in more depth, but the short version is the same as CSI: in-tree providers moved out-of-tree, and --cloud-provider=external is how a cluster says \u0026ldquo;ask a separate binary.\u0026rdquo;\nThat separate binary runs as a Deployment, leader-elected, watching the Kubernetes API and calling out to your infrastructure\u0026rsquo;s API:\ngraph TB subgraph CCM[\"Cloud Controller Manager - Deployment, leader-elected\"] NC[\"Node Controller\"] NLC[\"Node Lifecycle Controller\"] RC[\"Route Controller\"] SC[\"Service Controller\"] IF[\"cloudprovider.Interface\\n(your implementation)\"] NC --\u003e|Instances / InstancesV2| IF NLC --\u003e|Instances / InstancesV2| IF RC --\u003e|Routes| IF SC --\u003e|LoadBalancer| IF end K8S[(\"Kubernetes API Server\")] CLOUD[(\"Your Cloud API\\nVMs, Load Balancers, Networks\")] K8S \u003c--\u003e|watch / patch Nodes, Services, Routes| CCM IF --\u003e|create, attach, delete| CLOUD graph TB subgraph CCM[\"Cloud Controller Manager - Deployment, leader-elected\"] NC[\"Node Controller\"] NLC[\"Node Lifecycle Controller\"] RC[\"Route Controller\"] SC[\"Service Controller\"] IF[\"cloudprovider.Interface\\n(your implementation)\"] NC --\u003e|Instances / InstancesV2| IF NLC --\u003e|Instances / InstancesV2| IF RC --\u003e|Routes| IF SC --\u003e|LoadBalancer| IF end K8S[(\"Kubernetes API Server\")] CLOUD[(\"Your Cloud API\\nVMs, Load Balancers, Networks\")] K8S \u003c--\u003e|watch / patch Nodes, Services, Routes| CCM IF --\u003e|create, attach, delete| CLOUD ⤢ Everything downstream of that diagram is one interface and four controllers that call into it.\nThe cloudprovider.Interface Everything a CCM can do is exposed through one interface, k8s.io/cloud-provider\u0026rsquo;s Interface. You implement the pieces relevant to your infrastructure; each accessor also returns a bool so the controller manager can skip wiring up a controller you don\u0026rsquo;t support:\ntype Interface interface { Initialize(clientBuilder ControllerClientBuilder, stop \u0026lt;-chan struct{}) LoadBalancer() (LoadBalancer, bool) Instances() (Instances, bool) InstancesV2() (InstancesV2, bool) Zones() (Zones, bool) Clusters() (Clusters, bool) Routes() (Routes, bool) ProviderName() string HasClusterID() bool } LoadBalancer - GetLoadBalancer, EnsureLoadBalancer, UpdateLoadBalancer, EnsureLoadBalancerDeleted. Backs every type: LoadBalancer Service. Instances / InstancesV2 - node identity: provider ID, instance type, addresses, zone/region. InstancesV2 is the one to implement today - it\u0026rsquo;s a single batched InstanceMetadata call instead of five separate RPCs, and implementing it disables the Zones interface entirely (more on that in pitfalls). Zones - deprecated, superseded by the zone/region fields on InstancesV2.InstanceMetadata. Only implement it if you\u0026rsquo;re stuck on the legacy Instances interface. Routes - pod-network routing at the infrastructure level (cloud-native VPC routes instead of a CNI overlay). Most clusters running Calico, Cilium, or another overlay-capable CNI don\u0026rsquo;t need this at all. Clusters - vestigial; almost nothing implements it. Initialize is the interesting one: it hands you a ControllerClientBuilder and a stop channel, and you can spawn your own goroutines and controllers from there - the interface isn\u0026rsquo;t the only thing driving behavior, as we\u0026rsquo;ll see in the patterns section.\nThe Four Controllers Same shape as CSI\u0026rsquo;s sidecars: you don\u0026rsquo;t call your driver directly, the controllers do, in response to Kubernetes API state.\nController Watches Calls Node controller new Nodes Instances/InstancesV2 - labels addresses, zone, instance type; removes the uninitialized taint Node lifecycle controller Node Ready condition InstanceExists/InstanceShutdown - taints or deletes Nodes whose backing instance is gone or stopped Route controller Nodes, pod CIDRs Routes - creates/deletes cloud-native routes so pod traffic reaches the right node Service controller type: LoadBalancer Services LoadBalancer - provisions/updates/deletes the load balancer, patches status.loadBalancer.ingress All four are workqueue-based client-go controllers under the hood - same informer → workqueue → worker → syncHandler shape you\u0026rsquo;d write by hand for any custom controller.\nThe CCM as a Client, Not a Controller It\u0026rsquo;s easy to mistake \u0026ldquo;watches the Kubernetes API, reacts to changes\u0026rdquo; with \u0026ldquo;operator.\u0026rdquo; A CCM does watch and react, but the write side of that loop is much narrower than an Operator\u0026rsquo;s or a Cluster API infrastructure provider\u0026rsquo;s. An Operator (or a Cluster API provider) typically owns one or more CRDs: it defines the schema, runs a reconcile loop that drives observed state toward spec, and both reads and writes its own custom resources, sometimes provisioning infrastructure as a result. A CCM does none of that. It defines no CRDs, runs no admission webhook for custom types, and your own cloudprovider.Interface implementation never touches a kubeClient at all.\nLook back at the four-controllers table above: every Kubernetes write (patching a Node\u0026rsquo;s labels, removing a taint, deleting a Node, patching Service.status) happens inside the four controllers k8s.io/cloud-provider already ships, not in code you write. Your job is answering their questions (does this instance exist, what\u0026rsquo;s its zone, does this Service have a load balancer yet), and every answer you give is a call to somebody else\u0026rsquo;s API, not a write to Kubernetes:\ngraph LR subgraph Provided[\"k8s.io/cloud-provider - the only code that writes to Kubernetes\"] NC[Node Controller] NLC[Node Lifecycle Controller] RC[Route Controller] SC[Service Controller] end subgraph YourCode[\"Your code - implements cloudprovider.Interface only\"] IF[\"Instances / InstancesV2 /\\nLoadBalancer / Routes\"] end K8S[(\"Kubernetes API\\nNode, Service\")] CLOUD[(\"Your Cloud API\")] Provided \u003c--\u003e|patch/delete Node,\\npatch Service.status| K8S Provided --\u003e|ask questions, on a fixed period| YourCode YourCode --\u003e|answer, via HTTP/gRPC| CLOUD graph LR subgraph Provided[\"k8s.io/cloud-provider - the only code that writes to Kubernetes\"] NC[Node Controller] NLC[Node Lifecycle Controller] RC[Route Controller] SC[Service Controller] end subgraph YourCode[\"Your code - implements cloudprovider.Interface only\"] IF[\"Instances / InstancesV2 /\\nLoadBalancer / Routes\"] end K8S[(\"Kubernetes API\\nNode, Service\")] CLOUD[(\"Your Cloud API\")] Provided \u003c--\u003e|patch/delete Node,\\npatch Service.status| K8S Provided --\u003e|ask questions, on a fixed period| YourCode YourCode --\u003e|answer, via HTTP/gRPC| CLOUD ⤢ That narrow footprint is also why a CCM\u0026rsquo;s consistency model can afford to be \u0026ldquo;best-effort sync\u0026rdquo; instead of exactly-once. Each controller runs its own question on its own fixed period (MonitorNodes on nodeMonitorPeriod, service syncs pulled off a workqueue by N workers) instead of as one atomic transaction across Kubernetes and your cloud. A failed EnsureLoadBalancer this cycle isn\u0026rsquo;t a lost transaction that needs a saga to unwind; it\u0026rsquo;s a workqueue item that gets requeued and tried again next resync.\nThat\u0026rsquo;s a deliberate simplification. It\u0026rsquo;s why the sequential, hand-rolled orchestration from the reconcile-in-the-backend section is unnecessary, and it\u0026rsquo;s why the best-effort retry pattern in the concurrency section above is safe: \u0026ldquo;retry next cycle\u0026rdquo; is a far more forgiving contract than \u0026ldquo;must not partially apply.\u0026rdquo;\nStrip away the Kubernetes framing and what\u0026rsquo;s left is a plain client-integration problem: your code holds a long-lived connection to someone else\u0026rsquo;s API, on a fixed poll cadence, at whatever scale the cluster grows to. That\u0026rsquo;s the same set of problems any SDK client talking to someone else\u0026rsquo;s backend has to solve - and it\u0026rsquo;s worth treating them as first-class instead of incidental.\nBeing a Well-Behaved Client Deadlines everywhere, not just at the edge. Every cloudprovider.Interface method already gets a ctx from the controller calling it - the mistake is dropping it once you\u0026rsquo;re inside your own client, whether that\u0026rsquo;s an http.Client built without NewRequestWithContext or a goroutine that outlives the call that spawned it. A dropped deadline doesn\u0026rsquo;t just risk hanging - it leaves your backend holding a connection or a slot for work nobody is waiting on anymore:\nfunc (c *Client) createLoadBalancer(ctx context.Context, spec LBSpec) (*LB, error) { req, err := http.NewRequestWithContext(ctx, http.MethodPost, c.baseURL+\u0026#34;/load-balancers\u0026#34;, encode(spec)) if err != nil { return nil, err } resp, err := c.http.Do(req) // ctx.Done() aborts the round trip itself, not just the wait if err != nil { return nil, err } defer resp.Body.Close() return decodeLB(resp) } Classify errors before you retry them. Not every failure is worth retrying, and retrying the wrong ones burns the budget you need for the right ones. A 400 means your request is malformed - retrying it just repeats the mistake. A 429 or 503 means \u0026ldquo;try again later\u0026rdquo; - that\u0026rsquo;s the one your backoff is for:\ntype errClass int const ( errTerminal errClass = iota errRetryable ) func classify(err error) errClass { var apiErr *cloudapi.Error if errors.As(err, \u0026amp;apiErr) { switch apiErr.StatusCode { case http.StatusTooManyRequests, http.StatusServiceUnavailable, http.StatusGatewayTimeout: return errRetryable case http.StatusBadRequest, http.StatusNotFound, http.StatusConflict: return errTerminal } } if errors.Is(err, context.DeadlineExceeded) { return errRetryable } return errTerminal } Backoff needs jitter, or you\u0026rsquo;ve just rescheduled the thundering herd. Plain exponential backoff - base * 2^attempt for every caller - synchronizes retries instead of spreading them: every replica that failed at the same instant also retries at the same instant, at every subsequent step. This is the same failure mode written up in AWS\u0026rsquo;s Builders\u0026rsquo; Library work on backoff and jitter - the fix isn\u0026rsquo;t backoff, it\u0026rsquo;s randomizing within the backoff window so retries decorrelate:\nfunc backoff(attempt int) time.Duration { base, max := 250*time.Millisecond, 30*time.Second d := base * time.Duration(1\u0026lt;\u0026lt;attempt) if d \u0026gt; max { d = max } return time.Duration(rand.Int63n(int64(d))) // full jitter: uniform in [0, d) } Rate limit yourself - a bursty client is the backend\u0026rsquo;s reliability problem, not just yours. A CCM watching thousands of Nodes and Services can generate real bursts: a cold-start sync on restart, a mass Node replacement during a cluster upgrade, a Service migration. From the backend\u0026rsquo;s side, an unthrottled burst looks the same regardless of intent, and it\u0026rsquo;s exactly what turns a blip into a retry storm - each layer\u0026rsquo;s retries amplifying load on the layer below it. Pairing a client-side token bucket with the backend\u0026rsquo;s documented limits makes your throughput predictable instead of bursty-then-reactive:\ntype Client struct { http *http.Client limiter *rate.Limiter // e.g. rate.NewLimiter(rate.Limit(5), 10): 5 req/s, burst 10 } func (c *Client) do(ctx context.Context, req *http.Request) (*http.Response, error) { if err := c.limiter.Wait(ctx); err != nil { // blocks for a token, or dies with ctx return nil, err } return c.http.Do(req) } Honor the backend\u0026rsquo;s own backpressure signal instead of guessing. When you do get throttled, a 429 usually comes with a Retry-After header - the backend telling you, specifically, how long it needs. That\u0026rsquo;s better information than your own backoff curve; fold it in when present instead of overriding it:\nfunc retryAfter(resp *http.Response) (time.Duration, bool) { v := resp.Header.Get(\u0026#34;Retry-After\u0026#34;) if v == \u0026#34;\u0026#34; { return 0, false } if secs, err := strconv.Atoi(v); err == nil { return time.Duration(secs) * time.Second, true } if t, err := http.ParseTime(v); err == nil { return time.Until(t), true } return 0, false } Jitter your resync cadence too, not just your retries. The node lifecycle controller\u0026rsquo;s MonitorNodes loop and the service controller\u0026rsquo;s periodic resync both run on a fixed period. That\u0026rsquo;s fine for one replica - but a managed backend usually isn\u0026rsquo;t serving one CCM, it\u0026rsquo;s serving one per cluster, across a whole fleet. If every replica restarts around the same event - a rollout, a shared dependency recovering from an outage - and all of them resync on the same fixed period, they knock on the backend\u0026rsquo;s door in lockstep, forever. wait.JitterUntil, already in k8s.io/apimachinery, spreads that out for free:\n// Fixed period: every replica in the fleet converges on the same tick. wait.UntilWithContext(ctx, c.MonitorNodes, c.nodeMonitorPeriod) // Jittered: each replica\u0026#39;s period wobbles +/-20%, so a fleet restart doesn\u0026#39;t sync up. wait.JitterUntil(func() { c.MonitorNodes(ctx) }, c.nodeMonitorPeriod, 0.2, true, ctx.Done()) None of this is Kubernetes-specific - it\u0026rsquo;s the same checklist you\u0026rsquo;d run for any SDK client talking to someone else\u0026rsquo;s API. The Kubernetes-specific part is just knowing where your write boundary actually ends: at the edge of cloudprovider.Interface, not inside the four controllers, and not anywhere near a kubeClient.\nNode Initialization Every Node that joins a cluster running --cloud-provider=external is born with a taint:\nnode.cloudprovider.kubernetes.io/uninitialized:NoSchedule Nothing gets scheduled onto it until the CCM removes that taint - which it only does after successfully fetching instance metadata from your cloud:\nsequenceDiagram participant Kubelet participant K8s as Kubernetes API participant NC as Node Controller participant IF as cloudprovider.Interface participant Cloud as Your Cloud API Kubelet-\u003e\u003eK8s: register Node (tainted: uninitialized) K8s-\u003e\u003eNC: Node created event NC-\u003e\u003eIF: InstanceMetadata(ctx, node) IF-\u003e\u003eCloud: look up instance by providerID / name Cloud--\u003e\u003eIF: instance type, addresses, zone, region IF--\u003e\u003eNC: cloudprovider.InstanceMetadata{...} NC-\u003e\u003eK8s: patch Node - addresses, labels, providerID NC-\u003e\u003eK8s: remove uninitialized taint K8s--\u003e\u003eKubelet: Node schedulable sequenceDiagram participant Kubelet participant K8s as Kubernetes API participant NC as Node Controller participant IF as cloudprovider.Interface participant Cloud as Your Cloud API Kubelet-\u003e\u003eK8s: register Node (tainted: uninitialized) K8s-\u003e\u003eNC: Node created event NC-\u003e\u003eIF: InstanceMetadata(ctx, node) IF-\u003e\u003eCloud: look up instance by providerID / name Cloud--\u003e\u003eIF: instance type, addresses, zone, region IF--\u003e\u003eNC: cloudprovider.InstanceMetadata{...} NC-\u003e\u003eK8s: patch Node - addresses, labels, providerID NC-\u003e\u003eK8s: remove uninitialized taint K8s--\u003e\u003eKubelet: Node schedulable ⤢ The taint is the whole safety mechanism: if the CCM is down when a Node joins, that Node sits NoSchedule forever - no pods land on it, silently, until the CCM comes back and processes it. It\u0026rsquo;s also why the CCM should run with more than one replica behind leader election; a single-replica CCM restart during a scale-up event stalls every new Node until the pod comes back.\nNode Lifecycle: Shutdown vs. Deleted Not every cloud treats a stopped instance the same way. Some delete it outright; others leave it around, powered off, still holding its IP and disks. The node lifecycle controller has to handle both without mixing them up - deleting a Node object for an instance that\u0026rsquo;s merely stopped would evict every pod on it unnecessarily:\nsequenceDiagram participant NLC as Node Lifecycle Controller participant IF as cloudprovider.Interface participant Cloud as Your Cloud API participant K8s as Kubernetes API Note over NLC: Node.status.conditions[Ready] != True NLC-\u003e\u003eIF: InstanceExists(ctx, node) IF-\u003e\u003eCloud: look up instance alt instance gone entirely Cloud--\u003e\u003eIF: not found IF--\u003e\u003eNLC: false NLC-\u003e\u003eK8s: delete Node object else instance still exists Cloud--\u003e\u003eIF: found IF--\u003e\u003eNLC: true NLC-\u003e\u003eIF: InstanceShutdown(ctx, node) IF-\u003e\u003eCloud: check power state Cloud--\u003e\u003eIF: shutdown = true IF--\u003e\u003eNLC: true NLC-\u003e\u003eK8s: taint Node (out-of-service) end sequenceDiagram participant NLC as Node Lifecycle Controller participant IF as cloudprovider.Interface participant Cloud as Your Cloud API participant K8s as Kubernetes API Note over NLC: Node.status.conditions[Ready] != True NLC-\u003e\u003eIF: InstanceExists(ctx, node) IF-\u003e\u003eCloud: look up instance alt instance gone entirely Cloud--\u003e\u003eIF: not found IF--\u003e\u003eNLC: false NLC-\u003e\u003eK8s: delete Node object else instance still exists Cloud--\u003e\u003eIF: found IF--\u003e\u003eNLC: true NLC-\u003e\u003eIF: InstanceShutdown(ctx, node) IF-\u003e\u003eCloud: check power state Cloud--\u003e\u003eIF: shutdown = true IF--\u003e\u003eNLC: true NLC-\u003e\u003eK8s: taint Node (out-of-service) end ⤢ InstanceExists returning false triggers immediate Node deletion - get that check wrong (e.g. a transient API timeout misread as \u0026ldquo;not found\u0026rdquo;) and you\u0026rsquo;ll delete Nodes for instances that are still very much running.\nService → LoadBalancer Provisioning The one most people actually care about:\nsequenceDiagram actor User participant K8s as Kubernetes API participant SC as Service Controller participant LB as LoadBalancer (your impl) participant Cloud as Your Cloud API User-\u003e\u003eK8s: create Service{type: LoadBalancer} K8s-\u003e\u003eSC: Service event SC-\u003e\u003eLB: GetLoadBalancer(ctx, clusterName, svc) LB--\u003e\u003eSC: exists = false SC-\u003e\u003eLB: EnsureLoadBalancer(ctx, clusterName, svc, nodes) LB-\u003e\u003eCloud: create load balancer, ports from svc.Spec.Ports Cloud-\u003e\u003eCloud: provision, assign public IP Cloud--\u003e\u003eLB: LB ID, ingress IP LB-\u003e\u003eCloud: attach target nodes LB--\u003e\u003eSC: LoadBalancerStatus{Ingress: [{IP: ...}]} SC-\u003e\u003eK8s: patch Service.status.loadBalancer.ingress K8s--\u003e\u003eUser: kubectl get svc shows EXTERNAL-IP sequenceDiagram actor User participant K8s as Kubernetes API participant SC as Service Controller participant LB as LoadBalancer (your impl) participant Cloud as Your Cloud API User-\u003e\u003eK8s: create Service{type: LoadBalancer} K8s-\u003e\u003eSC: Service event SC-\u003e\u003eLB: GetLoadBalancer(ctx, clusterName, svc) LB--\u003e\u003eSC: exists = false SC-\u003e\u003eLB: EnsureLoadBalancer(ctx, clusterName, svc, nodes) LB-\u003e\u003eCloud: create load balancer, ports from svc.Spec.Ports Cloud-\u003e\u003eCloud: provision, assign public IP Cloud--\u003e\u003eLB: LB ID, ingress IP LB-\u003e\u003eCloud: attach target nodes LB--\u003e\u003eSC: LoadBalancerStatus{Ingress: [{IP: ...}]} SC-\u003e\u003eK8s: patch Service.status.loadBalancer.ingress K8s--\u003e\u003eUser: kubectl get svc shows EXTERNAL-IP ⤢ EnsureLoadBalancer is called on create and on every relevant update - a node joining or leaving, a port changing, an annotation flipping. It has to be safe to call repeatedly against the same desired state, which is the whole idempotency discussion in the pitfalls section below.\nRoute Reconciliation Only relevant if your CNI relies on cloud-native routing instead of an overlay:\nsequenceDiagram participant RC as Route Controller participant IF as cloudprovider.Interface participant Cloud as Your Cloud API RC-\u003e\u003eIF: ListRoutes(ctx, clusterName) IF-\u003e\u003eCloud: list routes on the cluster's network Cloud--\u003e\u003eIF: existing routes IF--\u003e\u003eRC: []cloudprovider.Route loop for each Node with a PodCIDR alt no matching route exists RC-\u003e\u003eIF: CreateRoute(ctx, clusterName, nameHint, route) IF-\u003e\u003eCloud: add route: PodCIDR -\u003e Node's internal IP end end loop for each stale route Note over RC: route's target Node no longer exists RC-\u003e\u003eIF: DeleteRoute(ctx, clusterName, route) IF-\u003e\u003eCloud: remove route end sequenceDiagram participant RC as Route Controller participant IF as cloudprovider.Interface participant Cloud as Your Cloud API RC-\u003e\u003eIF: ListRoutes(ctx, clusterName) IF-\u003e\u003eCloud: list routes on the cluster's network Cloud--\u003e\u003eIF: existing routes IF--\u003e\u003eRC: []cloudprovider.Route loop for each Node with a PodCIDR alt no matching route exists RC-\u003e\u003eIF: CreateRoute(ctx, clusterName, nameHint, route) IF-\u003e\u003eCloud: add route: PodCIDR -\u003e Node's internal IP end end loop for each stale route Note over RC: route's target Node no longer exists RC-\u003e\u003eIF: DeleteRoute(ctx, clusterName, route) IF-\u003e\u003eCloud: remove route end ⤢ Concurrency: Parallelize Carefully Even with a declarative backend, one part stays the CCM\u0026rsquo;s problem: attaching 50 nodes to a load balancer, or reconciling 200 Services, one at a time is slow, and the naive implementation of almost every reconcile loop is sequential by default - a for loop over targets, awaiting each call before starting the next. Fixing that has its own set of gotchas that are easy to get wrong in the other direction.\nTwo different knobs, don\u0026rsquo;t confuse them. The controllers themselves have a concurrency setting - --concurrent-service-syncs controls how many Services get reconciled in parallel across the workqueue, by spawning that many worker goroutines pulling off one shared queue (client-go\u0026rsquo;s workqueue guarantees no two workers ever get the same key at once):\n// Controller-level: N Services reconciled in parallel - real code from // k8s.io/cloud-provider\u0026#39;s service controller. for i := 0; i \u0026lt; workers; i++ { go wait.UntilWithContext(ctx, c.serviceWorker, time.Second) } That\u0026rsquo;s a separate knob from what happens inside one EnsureLoadBalancer call. Turning up --concurrent-service-syncs just gets you more reconciles stuck in the same slow sequential loop at once, unless the loop itself is fixed too:\n// Sequential by default - one slow or failed call blocks every node after it. func (l *loadBalancer) attachTargets(ctx context.Context, lb *LB, nodes []*v1.Node) error { for _, node := range nodes { if err := l.client.AttachTarget(ctx, lb.ID, node.Name); err != nil { return err } } return nil } Bound your concurrency, don\u0026rsquo;t remove it. Swapping that loop for go on every iteration just moves the bottleneck - you\u0026rsquo;ve traded one slow round trip at a time for a burst that trips the backend\u0026rsquo;s rate limiter. A semaphore-bounded errgroup, sized to what the backend can actually absorb concurrently, is the fix - not unbounded fan-out:\nfunc (l *loadBalancer) attachTargets(ctx context.Context, lb *LB, nodes []*v1.Node) error { g, ctx := errgroup.WithContext(ctx) g.SetLimit(8) // bounded to what the backend can absorb, not len(nodes) for _, node := range nodes { g.Go(func() error { return l.client.AttachTarget(ctx, lb.ID, node.Name) }) } return g.Wait() } Respect ordering dependencies. Attaching targets depends on the load balancer existing first - that edge can\u0026rsquo;t be parallelized away. Parallelize within a level of independent work, not across a dependency:\nfunc (l *loadBalancer) ensure(ctx context.Context, spec LBSpec) (*LB, error) { lb, err := l.client.CreateLoadBalancer(ctx, spec) // must happen first if err != nil { return nil, err } // targets are independent of each other, so this is safe to bound-parallelize if err := l.attachTargets(ctx, lb, spec.Targets); err != nil { return nil, err } return lb, nil } Shared state needs its own synchronization. A shared instance/TTL cache (see the patterns section below), or any counter/map touched from multiple reconcile goroutines, needs a mutex the moment more than one goroutine can touch it:\ntype instanceCache struct { mu sync.RWMutex byID map[string]*Instance } func (c *instanceCache) get(id string) (*Instance, bool) { c.mu.RLock() defer c.mu.RUnlock() inst, ok := c.byID[id] return inst, ok } func (c *instanceCache) set(id string, inst *Instance) { c.mu.Lock() defer c.mu.Unlock() c.byID[id] = inst // unguarded writes here are a data race the moment // two reconciles run concurrently } Partial failure needs per-item retry, and that shapes which primitive you reach for. errgroup.WithContext cancels every sibling goroutine the instant one fails - correct for \u0026ldquo;abort the whole operation on first failure,\u0026rdquo; wrong for \u0026ldquo;attempt all 50 attachments best-effort and report which ones failed.\u0026rdquo; The tempting alternative, rolling back what succeeded, is itself more API calls that can also fail. A plain sync.WaitGroup with per-item error collection lets the rest finish and hands back exactly what to retry next reconcile:\nfunc (l *loadBalancer) attachTargetsBestEffort(ctx context.Context, lb *LB, nodes []*v1.Node) []string { var ( wg sync.WaitGroup mu sync.Mutex failed []string sem = make(chan struct{}, 8) ) for _, node := range nodes { wg.Add(1) go func(name string) { defer wg.Done() sem \u0026lt;- struct{}{} defer func() { \u0026lt;-sem }() if err := l.client.AttachTarget(ctx, lb.ID, name); err != nil { mu.Lock() failed = append(failed, name) mu.Unlock() } }(node.Name) } wg.Wait() return failed // retried on the next reconcile - no rollback needed } Implementing in Go The interface methods themselves are mechanical - a main.go that imports your provider package for its init() side effect and hands off to app.NewCloudControllerManagerCommand, a Cloud struct whose accessors return (yourImpl, true) or (nil, false), and per-method translations between v1.Node/v1.Service and your cloud API\u0026rsquo;s request shapes. None of that is where the interesting decisions live - kubernetes/cloud-provider\u0026rsquo;s sample package is a complete, working skeleton of exactly that scaffolding, worth copying wholesale rather than retyping. One binary, no controller/node split like CSI has, since everything a CCM does is an API call, never a local mount:\ncmd/ main.go # wires the provider factory + calls app.NewCloudControllerManagerCommand internal/ cloud/ cloud.go # implements cloudprovider.Interface instances.go # InstancesV2 loadbalancer.go # LoadBalancer routes.go # Routes client.go # your cloud API client The part actually worth designing carefully is what loadbalancer.go and client.go do when a Service wants a load balancer - which is the next two sections.\nReconcile in the Backend, Not in the CCM If you\u0026rsquo;re building the CCM and the infrastructure API behind it - a home lab, a private cloud, an internal platform - this is the highest-leverage decision you\u0026rsquo;ll make, and it\u0026rsquo;s easy to get backwards.\nThe naive shape: your infrastructure models a load balancer as several separate resources - a front-end, a target group or pool, pool members, a health check, a listener. EnsureLoadBalancer walks that graph itself: create the target group, wait for it, create each member one call at a time, wait, create the listener, attach the health check, poll until everything reports healthy. Every one of those is a round trip the CCM is now personally responsible for sequencing, retrying, and recovering from a partial failure on.\nsequenceDiagram participant CCM as CCM (naive) participant API as Backend API Note over CCM,API: Client-side orchestration - CCM walks the resource graph CCM-\u003e\u003eAPI: create target group API--\u003e\u003eCCM: group ID CCM-\u003e\u003eAPI: add member 1 CCM-\u003e\u003eAPI: add member 2 CCM-\u003e\u003eAPI: add member N CCM-\u003e\u003eAPI: create listener CCM-\u003e\u003eAPI: attach health check loop poll until healthy CCM-\u003e\u003eAPI: GET status end sequenceDiagram participant CCM as CCM (naive) participant API as Backend API Note over CCM,API: Client-side orchestration - CCM walks the resource graph CCM-\u003e\u003eAPI: create target group API--\u003e\u003eCCM: group ID CCM-\u003e\u003eAPI: add member 1 CCM-\u003e\u003eAPI: add member 2 CCM-\u003e\u003eAPI: add member N CCM-\u003e\u003eAPI: create listener CCM-\u003e\u003eAPI: attach health check loop poll until healthy CCM-\u003e\u003eAPI: GET status end ⤢ sequenceDiagram participant CCM as CCM (declarative) participant API as Backend API participant R as Backend's own reconciler Note over CCM,API: Server-side orchestration - CCM submits desired state once CCM-\u003e\u003eAPI: PUT desired state (ports, targets, health check - one spec) API-\u003e\u003eR: enqueue reconcile API--\u003e\u003eCCM: 202 Accepted, status=Pending R-\u003e\u003eR: create group, members, listener, health check loop poll or watch CCM-\u003e\u003eAPI: GET status API--\u003e\u003eCCM: Pending end API--\u003e\u003eCCM: Ready, ingress IP sequenceDiagram participant CCM as CCM (declarative) participant API as Backend API participant R as Backend's own reconciler Note over CCM,API: Server-side orchestration - CCM submits desired state once CCM-\u003e\u003eAPI: PUT desired state (ports, targets, health check - one spec) API-\u003e\u003eR: enqueue reconcile API--\u003e\u003eCCM: 202 Accepted, status=Pending R-\u003e\u003eR: create group, members, listener, health check loop poll or watch CCM-\u003e\u003eAPI: GET status API--\u003e\u003eCCM: Pending end API--\u003e\u003eCCM: Ready, ingress IP ⤢ // Naive: the CCM walks the resource graph itself, one call per node. func (l *loadBalancer) ensureNaive(ctx context.Context, spec LBSpec) (string, error) { group, err := l.client.CreateTargetGroup(ctx, spec.Name) if err != nil { return \u0026#34;\u0026#34;, fmt.Errorf(\u0026#34;creating target group: %w\u0026#34;, err) } for _, target := range spec.Targets { if err := l.client.AddGroupMember(ctx, group.ID, target); err != nil { // group now exists half-populated - every future call has to // detect and repair this from the Kubernetes side return \u0026#34;\u0026#34;, fmt.Errorf(\u0026#34;adding member %s: %w\u0026#34;, target, err) } } listener, err := l.client.CreateListener(ctx, group.ID, spec.Ports) if err != nil { return \u0026#34;\u0026#34;, fmt.Errorf(\u0026#34;creating listener: %w\u0026#34;, err) } return l.pollUntilHealthy(ctx, listener.ID) } // Declarative: the CCM submits desired state once, the backend owns the graph. func (l *loadBalancer) ensureDeclarative(ctx context.Context, spec LBSpec) (string, error) { op, err := l.client.SubmitDesiredState(ctx, spec) // one PUT; backend reconciles async if err != nil { return \u0026#34;\u0026#34;, fmt.Errorf(\u0026#34;submitting desired state: %w\u0026#34;, err) } return l.pollUntilReady(ctx, op.ID) // same polling shape, far less to get wrong } The second shape is exactly what Kubernetes itself does to you as a client - you POST a Deployment, you don\u0026rsquo;t walk the apiserver through creating a ReplicaSet, then Pods, then container statuses one call at a time. The backend API should own its own convergence loop, its own idempotency, and its own retry logic, behind one declarative endpoint. The CCM\u0026rsquo;s job shrinks to submitting desired state and translating the eventual status back into v1.LoadBalancerStatus - which is also just a much smaller surface for bugs to hide in.\nIf you\u0026rsquo;re targeting a third-party cloud\u0026rsquo;s API, you don\u0026rsquo;t get this choice - AWS, Hetzner, and DigitalOcean\u0026rsquo;s APIs are what they are, and orchestrating their sub-resources from the CCM is simply required. But if you\u0026rsquo;re the one writing both sides, don\u0026rsquo;t reinvent client-side orchestration that a declarative backend endpoint would make disappear.\nDeployment A CCM is a single Deployment, not a DaemonSet - there\u0026rsquo;s no per-node privileged work, everything routes through the cloud API:\napiVersion: apps/v1 kind: Deployment metadata: name: cloud-controller-manager namespace: kube-system spec: replicas: 2 selector: matchLabels: app: cloud-controller-manager template: metadata: labels: app: cloud-controller-manager spec: serviceAccountName: cloud-controller-manager tolerations: - key: node.cloudprovider.kubernetes.io/uninitialized value: \u0026#34;true\u0026#34; effect: NoSchedule - key: node-role.kubernetes.io/control-plane effect: NoSchedule containers: - name: cloud-controller-manager image: your-registry/cloud-controller-manager:latest args: - --cloud-provider=cloud.example.com - --leader-elect=true - --use-service-account-credentials=true - --controllers=cloud-node,cloud-node-lifecycle,service,route Two things worth calling out:\nThe CCM itself must tolerate the taint it\u0026rsquo;s responsible for removing. Skip the toleration and it can never schedule in the first place - a bootstrapping deadlock. replicas: 2 with leader election, not 1 - the node-initialization deadlock above is the reason. RBAC needs get/list/watch/patch/update on nodes, services, endpoints, and events, plus create/update on nodes/status and services/status. leases.coordination.k8s.io for leader election.\nPatterns from Production CCMs Once the four controllers and the basic interface implementation work, the interesting decisions are in the details. A few worth stealing, pulled from Hetzner\u0026rsquo;s and DigitalOcean\u0026rsquo;s CCMs:\nRate-limit circuit breaker. Wrap your cloud API client so a 429/rate-limit response sets a local \u0026ldquo;exceeded until T\u0026rdquo; flag, and short-circuit every subsequent call against that flag until it expires - instead of hammering an already-rate-limited API with retries from every controller simultaneously.\nTTL cache with per-caller override. A single instance cache backs Instances, Routes, and LoadBalancer lookups, with a short default TTL - but a caller that only needs a slow-changing field (e.g. the routes controller reading a network attachment) can explicitly request a looser TTL and skip an API call it doesn\u0026rsquo;t need freshness for. One cache, tunable per read pattern, instead of either always-fresh or a single global TTL that\u0026rsquo;s wrong for someone.\nField-by-field declarative reconcile. Instead of one function that mutates a load balancer however it sees fit, decompose it: changeAlgorithm, changeType, attachToNetwork, togglePublicInterface - each diffs one attribute against desired state, returns whether it changed anything, aggregated into \u0026ldquo;does the caller need to re-fetch the object.\u0026rdquo; Each piece is independently testable and independently readable.\nA second controller outside the interface. Initialize() hands you a client builder specifically so you can run additional controllers the cloudprovider.Interface has no hook for - DigitalOcean runs a standalone workqueue controller reconciling a shared cluster firewall, completely separate from the service controller. The interface covers the common cases; it doesn\u0026rsquo;t limit what else your CCM can manage.\nAdmission webhook validating against the real API. Rather than let a malformed LB annotation fail asynchronously three reconciles deep in EnsureLoadBalancer, DigitalOcean runs a ValidatingWebhookConfiguration that builds the actual load-balancer request from the Service\u0026rsquo;s annotations and dry-validates it against the cloud API at admission time - a bad kubectl apply gets rejected synchronously instead of silently failing in a controller loop nobody\u0026rsquo;s watching.\nPeriodic resync as a second reconciliation source. Event-driven reconciliation misses things - a webhook that didn\u0026rsquo;t fire, an informer resync that raced a restart. A ticker-driven background sync (re-tagging resources with ownership metadata, in DigitalOcean\u0026rsquo;s case) running independently of the event-driven path catches what the watch-based controllers miss.\nOwnership guards before mutating. Before deleting or modifying any cloud resource, check that it\u0026rsquo;s actually tagged/labeled as owned by this cluster. Skipping this check is how a CCM ends up deleting a load balancer a human created by hand with the same name.\nMigrating Off an In-Tree Provider If you\u0026rsquo;re moving a cluster from a built-in cloud provider to an external CCM, the sequencing matters more than the CCM code itself:\nSet --cloud-provider=external on the kubelet and kube-apiserver/kube-controller-manager - this disables the in-tree provider and starts tainting new Nodes uninitialized. Deploy the CCM before any Node restarts, or joins with the new flag - otherwise Nodes sit tainted with nothing to untaint them. Never run the in-tree provider and an external CCM against the same cluster simultaneously. Both will try to reconcile the same Services and routes, and you\u0026rsquo;ll get flapping load balancer state as they fight over target lists. Existing Nodes don\u0026rsquo;t automatically get the taint retroactively - only newly-registered Nodes do. A rolling node replacement (not an in-place kubelet flag flip) is the safer migration path for that reason. Common Pitfalls The monolithic EnsureLoadBalancer. The version of this function that first works tends to be one large function doing everything inline - validate, diff, create, attach, set health checks, handle TLS, in one pass. It works, until it doesn\u0026rsquo;t: one bug in TLS handling now blocks unrelated port changes from ever reconciling, and there\u0026rsquo;s no way to unit test \u0026ldquo;does the algorithm change logic work\u0026rdquo; without exercising the whole thing end to end. Decompose it field-by-field from the start, the way the patterns section above describes - it costs more lines up front and saves the debugging session later.\nInstancesV2 and Zones are mutually exclusive. The doc comment says it plainly - implementing InstancesV2 disables calls to Zones entirely, regardless of whether you also return true from Zones(). If zone/region data isn\u0026rsquo;t showing up on Nodes, check that it\u0026rsquo;s actually being set on InstanceMetadata.Zone/.Region, not in a Zones implementation nobody\u0026rsquo;s calling anymore.\nEnsureLoadBalancer idempotency. It\u0026rsquo;s called on every relevant Service or Node change, not just creation. Look up by a stable identifier (tag, not name - names get reused) before creating, and treat \u0026ldquo;already exists with the right config\u0026rdquo; as success, not an error.\nRoute controller you don\u0026rsquo;t need. If your CNI (Calico in overlay mode, Cilium, Flannel) already handles pod-to-pod routing, implementing Routes is pure overhead - it\u0026rsquo;s for clusters relying on cloud-native VPC routing instead of an overlay. Check what your CNI actually needs before building it.\nShutdown vs. deleted, mixed up. Deleting a Node object for an instance that\u0026rsquo;s merely powered off evicts every pod on it needlessly and forces a full reschedule. InstanceExists and InstanceShutdown are two different questions - mixing them up is the single easiest way to turn a planned maintenance window into an unplanned one.\nTesting Unit test against a fake cloudprovider.Interface implementation with in-memory state instead of a real cloud client - the same role k8s.io/mount-utils\u0026rsquo;s FakeMounter plays for CSI node testing:\nfunc TestEnsureLoadBalancer_Idempotent(t *testing.T) { lb := newLoadBalancer(fakeClient) svc := newService(\u0026#34;my-svc\u0026#34;, corev1.ServiceTypeLoadBalancer) status1, err := lb.EnsureLoadBalancer(ctx, \u0026#34;cluster\u0026#34;, svc, nodes) require.NoError(t, err) status2, err := lb.EnsureLoadBalancer(ctx, \u0026#34;cluster\u0026#34;, svc, nodes) require.NoError(t, err) assert.Equal(t, status1.Ingress, status2.Ingress) assert.Equal(t, 1, fakeClient.CreateLoadBalancerCallCount()) } For the controllers themselves - node, node lifecycle, route, service - envtest against a real kube-apiserver binary with your fake cloud plugged in catches the class of bug unit tests miss: taint ordering, informer resync races, and RBAC gaps that only show up against real API server validation.\nWhat You End Up With A single Deployment implementing cloudprovider.Interface against your infrastructure\u0026rsquo;s API Four controllers you didn\u0026rsquo;t write, driving your implementation: node, node lifecycle, route, service Nodes that get labeled, addressed, and untainted automatically as they join type: LoadBalancer Services that provision real infrastructure and get a real IP back Optionally, pod-network routes reconciled at the cloud level instead of by your CNI The interface is small - six methods, most of which you\u0026rsquo;ll return (nil, false) from. The actual complexity is entirely in the reconciliation logic behind the two or three you do implement: idempotency under repeated calls, distinguishing \u0026ldquo;gone\u0026rdquo; from \u0026ldquo;stopped,\u0026rdquo; and not letting one function grow to do everything at once. kubernetes/cloud-provider is the right starting point - the sample package in that repo and the Hetzner/DigitalOcean CCMs linked above are worth having open while you build.\n","permalink":"https://syndbg.dev/posts/2026-08-06-writing-a-kubernetes-cloud-controller-manager/","summary":"\u003cp\u003eA Cloud Controller Manager is a bridge. On one side: a Kubernetes cluster that thinks in Nodes, Services, and Routes. On the other: whatever actually runs your machines and network - a hyperscaler like AWS, GCP, or Azure, or just as often a private cloud, a colocated rack, or a home lab. The CCM\u0026rsquo;s whole job is translation - when a Node joins, ask the infrastructure what it knows about that machine and translate it onto the Node object; when a \u003ccode\u003eService\u003c/code\u003e wants a load balancer, go create one and wire the ingress IP back into \u003ccode\u003estatus\u003c/code\u003e. Without it, \u003ccode\u003ekubectl get nodes\u003c/code\u003e never gets a region label and \u003ccode\u003etype: LoadBalancer\u003c/code\u003e never gets an IP - Kubernetes has no idea any of that infrastructure exists.\u003c/p\u003e","title":"Writing a Kubernetes Cloud Controller Manager: Bridging Your Infra to Kubernetes"},{"content":"What Is a Coding Harness? Agent = Model + Harness. The model is the reasoning engine. The harness is everything else — the guides that steer it before it acts, the sensors that catch it after it does. (Harness Engineering).\nIn practice a coding harness has several moving parts:\nContext injection — what the agent knows before it writes a line. CLAUDE.md files, system prompts, project conventions. Codebase indexing — a queryable map of the repo. Symbols, call graphs, semantic embeddings. Tool layer — what the agent can do. Read files, run tests, call linters, search the index. Feedback loop — sensors that observe output quality. Type errors, failing tests, lint violations, architecture drift. Memory / sync — persistence across sessions. Hooks that re-index on session start, sync on file write. graph TD A[Developer Intent] --\u003e B[Context InjectionCLAUDE.md · system prompt · conventions] B --\u003e C[Model] C --\u003e D[Tool Layerread · write · run tests · search index] D --\u003e E[Codebase Indexsymbols · embeddings · call graph] E --\u003e|semantic search results| C D --\u003e F[Feedback Looptype errors · lint · test failures] F --\u003e|sensor output| C C --\u003e G[Output] G --\u003e H[Memory / SyncPostToolUse hook · incremental re-index] H --\u003e E graph TD A[Developer Intent] --\u003e B[Context InjectionCLAUDE.md · system prompt · conventions] B --\u003e C[Model] C --\u003e D[Tool Layerread · write · run tests · search index] D --\u003e E[Codebase Indexsymbols · embeddings · call graph] E --\u003e|semantic search results| C D --\u003e F[Feedback Looptype errors · lint · test failures] F --\u003e|sensor output| C C --\u003e G[Output] G --\u003e H[Memory / SyncPostToolUse hook · incremental re-index] H --\u003e E ⤢ The vocabulary here is not new. Russell \u0026amp; Norvig’s AI: A Modern Approach (1995) defined an agent as anything that perceives its environment through sensors and acts upon it through actuators. Chip Huyen applies this directly to modern LLM agents in Agents (2025): the model is the brain, tools are the actuators, observations are the sensors. Böckeler maps the same structure onto coding harnesses — guides are actuators (steer behavior forward), sensors are feedback mechanisms (observe what came out).\nThe naming changed. The pattern did not.\nMost developers using Claude Code today have the tool layer (built-in) and partial context injection (CLAUDE.md). Almost nobody has the indexing layer wired up. That gap is where the performance falls off a cliff.\nClaude Code: Hooks and Harness Points Claude Code exposes four hook events:\nHook Fires when Harness use SessionStart Agent session begins Re-index codebase, load project state PreToolUse Before any tool call Gate dangerous ops, inject context PostToolUse After Write/Edit Sync modified file back to index Stop Agent finishes Run validators, emit summaries These hooks are the harness attachment points. A SessionStart hook that runs incremental re-indexing means the agent starts every session with a fresh semantic map of the codebase. A PostToolUse hook on Write/Edit means the index never drifts more than one file behind reality.\nWithout these wired up, the agent is flying blind every session. It compensates by grepping. A lot.\nHooks also solve a problem that surfaces when running parallel workstreams with git worktrees. Each worktree is an isolated working directory — useful for running two agent sessions on different branches simultaneously without conflicts. But secrets and environment files (.env, credentials, tokens) do not copy across automatically. A SessionStart hook scoped to the worktree can detect the working directory, locate the canonical secrets from a shared location, and symlink or inject them without duplicating sensitive files across every checkout. The harness manages the plumbing so neither the developer nor the agent has to think about it.\nE.g\n#!/usr/bin/env bash # .claude/hooks/session-start.sh # Symlink .env from the main worktree into the current worktree on session start. MAIN_WORKTREE=$(git worktree list --porcelain | awk \u0026#39;NR==1{print $2}\u0026#39;) CURRENT_DIR=$(pwd) if [ \u0026#34;$CURRENT_DIR\u0026#34; != \u0026#34;$MAIN_WORKTREE\u0026#34; ] \u0026amp;\u0026amp; [ -f \u0026#34;$MAIN_WORKTREE/.env\u0026#34; ]; then ln -sf \u0026#34;$MAIN_WORKTREE/.env\u0026#34; \u0026#34;$CURRENT_DIR/.env\u0026#34; echo \u0026#34;Linked .env from $MAIN_WORKTREE\u0026#34; fi Wire it in .claude/settings.json:\n{ \u0026#34;hooks\u0026#34;: { \u0026#34;SessionStart\u0026#34;: [ { \u0026#34;matcher\u0026#34;: \u0026#34;\u0026#34;, \u0026#34;hooks\u0026#34;: [ { \u0026#34;type\u0026#34;: \u0026#34;command\u0026#34;, \u0026#34;command\u0026#34;: \u0026#34;bash .claude/hooks/session-start.sh\u0026#34; } ] } ] } } Runs once per session. If the working directory is a worktree (not main checkout), symlinks .env from main. No copies, no stale secrets, no manual setup per branch.\nOr, you can use Claude code with --worktree which also invokes the WorktreeCreate hook, as per the example in Claude\u0026rsquo;s docs.\nCursor: Where Codebase Indexing Comes Built-In Cursor does not make you wire this up yourself. It ships with persistent, background codebase indexing as a first-class feature. On session open, it already knows your symbols, your imports, your function signatures. Retrieval is a lookup, not an exploration.\nThis creates a felt difference in terms of accuracy, speed and cost-efficiency. Cursor doesn\u0026rsquo;t, based on my experience, significantly regress to greping and bisecting for repetitive work. This cuts token usage quite significantly.\nHow Cursor\u0026rsquo;s indexing pipeline (docs) works:\ngraph LR A[Files on disk] --\u003e B[AST chunkertree-sitter] B --\u003e C[Embedding modelOpenAI / custom] C --\u003e D[Turbopufferremote vector DB] D --\u003e|nearest-neighbor search| E[Query embedding] E --\u003e F[LLM context] subgraph Sync G[Merkle tree hash] --\u003e|simhash| H[Server: reuse existing index?] H --\u003e|92% similarity across org clones| I[Skip re-embedding] end A --\u003e G graph LR A[Files on disk] --\u003e B[AST chunkertree-sitter] B --\u003e C[Embedding modelOpenAI / custom] C --\u003e D[Turbopufferremote vector DB] D --\u003e|nearest-neighbor search| E[Query embedding] E --\u003e F[LLM context] subgraph Sync G[Merkle tree hash] --\u003e|simhash| H[Server: reuse existing index?] H --\u003e|92% similarity across org clones| I[Skip re-embedding] end A --\u003e G ⤢ Cursor indexes without storing filenames or source code — filenames are obfuscated, chunks are encrypted. Content proofs verify the client holds the file before results are returned.\nThe Merkle tree + simhash combination is what makes org-wide index reuse possible. A Merkle tree hashes file content bottom-up — any change propagates to the root, so the root hash changes. Simhash produces a fingerprint of a document set where similar sets produce similar hashes. Together they let Cursor detect that your clone of a repo is 92% identical to one already indexed on the server, and skip re-embedding the matching chunks entirely.\ngraph TD R[\"Root hash\\nsha256(H_AB + H_CD)\"] H_AB[\"H_AB\\nsha256(H_A + H_B)\"] H_CD[\"H_CD\\nsha256(H_C + H_D)\"] H_A[\"H_A\\nsha256(file_a.go)\"] H_B[\"H_B\\nsha256(file_b.go)\"] H_C[\"H_C\\nsha256(file_c.go)\"] H_D[\"H_D ⚠️\\nsha256(file_d.go) CHANGED\"] R --\u003e H_AB R --\u003e H_CD H_AB --\u003e H_A H_AB --\u003e H_B H_CD --\u003e H_C H_CD --\u003e H_D graph TD R[\"Root hash\\nsha256(H_AB + H_CD)\"] H_AB[\"H_AB\\nsha256(H_A + H_B)\"] H_CD[\"H_CD\\nsha256(H_C + H_D)\"] H_A[\"H_A\\nsha256(file_a.go)\"] H_B[\"H_B\\nsha256(file_b.go)\"] H_C[\"H_C\\nsha256(file_c.go)\"] H_D[\"H_D ⚠️\\nsha256(file_d.go) CHANGED\"] R --\u003e H_AB R --\u003e H_CD H_AB --\u003e H_A H_AB --\u003e H_B H_CD --\u003e H_C H_CD --\u003e H_D ⤢ One file changes → its leaf hash changes → H_CD changes → root changes. Cursor walks the diff, re-embeds only changed leaves. Unchanged subtrees reuse the existing index via copy_from_namespace.\nMedian time-to-first-query for large repos dropped from 7.87s → 525ms after Merkle-based index reuse shipped. (Cursor: Secure Codebase Indexing) (note: for some reason the link only redirects to the CN article, not the English one.)\nThe vector store behind this is Turbopuffer. Cursor runs one namespace per codebase — active ones stay hot in memory/NVMe, inactive ones spill to object storage. At scale: 1T+ documents across 80M+ namespaces, 10GB/s peak ingestion. Cold namespaces resume without re-embedding via copy_from_namespace, which is how the 92% org-clone similarity translates into actual latency savings rather than just a stat. Cursor moved to Turbopuffer in November 2023 and cut semantic search costs by 20x. Agent accuracy improved up to 23.5% after the switch.\nFrom entire.io\u0026rsquo;s analysis: nearly 49% of all agent tool calls are search operations. That is the exploration tax — paid in latency and tokens on every task that Cursor\u0026rsquo;s index would answer instantly.\nAgent tool call breakdown. Nearly half of all calls are search. Source: entire.io — Improving Agentic Search in Coding Agents\nTool execution is only 0.4% of wall-clock time. The bottleneck is model inference and planning, not search speed. Source: entire.io — Improving Agentic Search in Coding Agents\nLast but not least, Claude Code on the other hand has the more pluggable system, but completely lacks in this department, resorting to third-party tools we\u0026rsquo;ll see in a bit.\nWe Have Done This Before Here is the part that should sting: none of this is new.\nctags (1978) built symbol indexes so editors could jump to definitions without reading every file. cscope (1985) did cross-reference search across C codebases. Language servers (LSP, 2016) gave every editor a real-time semantic model of the code — go-to-definition, find-references, hover docs — without re-parsing anything.\nThe entire Language Server Protocol exists because editors kept re-implementing the same code intelligence in isolation. Microsoft proposed a standard. The ecosystem converged. Problem solved for the editor layer.\nNow in 2026 we are building the same thing again for AI agents. Semantic codebase indexes. Symbol-aware chunking. Incremental sync. The concepts are identical. The target consumer changed from \u0026ldquo;editor plugin\u0026rdquo; to \u0026ldquo;LLM tool call.\u0026rdquo;\nWe are not solving a new problem. We are re-plumbing an old one.\ntimeline title Code Intelligence: Same Problem, Different Consumer 1978 : ctags : Symbol index for editors : Jump-to-definition without reading every file 1985 : cscope : Cross-reference search across C codebases 2016 : LSP : Language Server Protocol : One standard, every editor : go-to-definition · find-references · hover docs 2024 : Cursor codebase indexing : AST chunking · embeddings · Turbopuffer : Pre-computed context for LLM queries 2025+ : Third-party tools, nothing established : Same pattern, new consumer : MCP tools instead of editor APIs timeline title Code Intelligence: Same Problem, Different Consumer 1978 : ctags : Symbol index for editors : Jump-to-definition without reading every file 1985 : cscope : Cross-reference search across C codebases 2016 : LSP : Language Server Protocol : One standard, every editor : go-to-definition · find-references · hover docs 2024 : Cursor codebase indexing : AST chunking · embeddings · Turbopuffer : Pre-computed context for LLM queries 2025+ : Third-party tools, nothing established : Same pattern, new consumer : MCP tools instead of editor APIs ⤢ Tools That Already Exist Here are some of the tools I\u0026rsquo;ve tried over the last few months.\nory/lumen Lumen is the closest thing the ecosystem has to a production-grade answer for Claude Code indexing. Single static binary, no Docker, no Python environment to manage. Install as a Claude Code plugin:\n/plugin install lumen@ory What it does well:\nAST-aware chunking via tree-sitter — splits at function boundaries, not arbitrary character windows. A retrieved chunk is a complete function, not lines 47–89 of one. An AST (Abstract Syntax Tree) is the parsed structure of source code — a tree where each node is a language construct: function, class, method, variable. Splitting at those boundaries means every retrieved chunk is a complete, meaningful unit.\ngraph TD A[\"source file\"] --\u003e B[\"module\"] B --\u003e C[\"class Config\"] B --\u003e D[\"function parse_config\"] B --\u003e E[\"function validate\"] D --\u003e F[\"param: path\"] D --\u003e G[\"body: open · load · validate · return\"] C --\u003e H[\"field: host\"] C --\u003e I[\"field: port\"] graph TD A[\"source file\"] --\u003e B[\"module\"] B --\u003e C[\"class Config\"] B --\u003e D[\"function parse_config\"] B --\u003e E[\"function validate\"] D --\u003e F[\"param: path\"] D --\u003e G[\"body: open · load · validate · return\"] C --\u003e H[\"field: host\"] C --\u003e I[\"field: port\"] ⤢ AST of a small Python module. Each node is a language construct. Chunking at function or class nodes produces self-contained units — not mid-expression fragments.\nNaive window chunking (500-char sliding window) might split this Python function arbitrarily:\n# Chunk 1 (chars 0-500) — cuts mid-function def parse_config(path: str) -\u0026gt; Config: \u0026#34;\u0026#34;\u0026#34;Load and validate config from disk.\u0026#34;\u0026#34;\u0026#34; with open(path) as f: raw = yaml.safe_load(f) if \u0026#34;database\u0026#34; not in raw: raise ValueError(\u0026#34;missing database # Chunk 2 (chars 500-1000) — starts mid-expression, no context key\u0026#34;) return Config( host=raw[\u0026#34;database\u0026#34;][\u0026#34;host\u0026#34;], port=raw[\u0026#34;database\u0026#34;].get(\u0026#34;port\u0026#34;, 5432), ) AST-aware chunking (tree-sitter, function boundary):\n# Chunk: parse_config — complete, self-contained def parse_config(path: str) -\u0026gt; Config: \u0026#34;\u0026#34;\u0026#34;Load and validate config from disk.\u0026#34;\u0026#34;\u0026#34; with open(path) as f: raw = yaml.safe_load(f) if \u0026#34;database\u0026#34; not in raw: raise ValueError(\u0026#34;missing database key\u0026#34;) return Config( host=raw[\u0026#34;database\u0026#34;][\u0026#34;host\u0026#34;], port=raw[\u0026#34;database\u0026#34;].get(\u0026#34;port\u0026#34;, 5432), ) The second chunk is retrievable, understandable, and embeds with full semantic context. The first produces garbage embeddings for the second half.\nCode-optimized embeddings: jina-embeddings-v2-base-code (512-dim), trained on code rather than general text.\nSQLite-vec as the vector store — embedded, zero infrastructure, single file on disk.\nSessionStart hook auto-wired on install. Index is fresh at the start of every session.\nBenchmarked: 26–39% token cost reduction, 28–53% faster sessions vs. baseline Claude Code.\nIn practice, indexing a medium Go service (220 files) takes ~9 minutes. A large polyglot monorepo (6,000+ files) can run for hours — and if you interrupt it, that nested repo logs zero chunks:\n$ lumen index . INFO Indexing nested repo /work/common-packages Embedded 36 chunks so far [12/12] ████████████████████████████████████████ 100% Indexing complete: 12 files, 36 chunks in 10.453s. INFO Indexing nested repo /work/project-a Embedded 2178 chunks so far [220/220] ████████████████████████████████████ 100% Indexing complete: 220 files, 2178 chunks in 9m9.039s. INFO Indexing nested repo /work/platform-repo Processing file 368/6375: environments/development/... [0367/6375] 6% Processing file 1601/6375: environments/production/... [1600/6375] 25% Processing file 3298/6375: environments/production/.../... [3297/6375] 52% ^C Nested repo /work/platform-repo: 0 files, 0 chunks in 3h15m16.412s. INFO Indexing nested repo /work/project-b Indexing complete: 176 files, 2168 chunks in 9m17.342s. Overall, I\u0026rsquo;ll be running this for hours and I am, when I want to capture the full scope.\nWhere it struggles:\nSQLite-vec works well for a single focused codebase. For local, single-machine use it is a reasonable embedded choice — zero infrastructure, single file. But the HNSW implementation it provides is a single-node, in-process index. No concurrent writers, no on-disk persistence separate from the SQLite file, no incremental segment merges.\nHNSW (Hierarchical Navigable Small World) is the dominant algorithm for approximate nearest-neighbor vector search — it builds a multi-layer graph where each layer skips further ahead, letting queries find close vectors in O(log n) rather than scanning everything.\ngraph TD subgraph \"Layer 2 (coarse)\" L2A((A)) --- L2E((E)) end subgraph \"Layer 1\" L1A((A)) --- L1C((C)) L1C --- L1E((E)) end subgraph \"Layer 0 (all nodes)\" L0A((A)) --- L0B((B)) L0B --- L0C((C)) L0C --- L0D((D)) L0D --- L0E((E)) end L2A -.-\u003e L1A L2E -.-\u003e L1E L1C -.-\u003e L0C graph TD subgraph \"Layer 2 (coarse)\" L2A((A)) --- L2E((E)) end subgraph \"Layer 1\" L1A((A)) --- L1C((C)) L1C --- L1E((E)) end subgraph \"Layer 0 (all nodes)\" L0A((A)) --- L0B((B)) L0B --- L0C((C)) L0C --- L0D((D)) L0D --- L0E((E)) end L2A -.-\u003e L1A L2E -.-\u003e L1E L1C -.-\u003e L0C ⤢ HNSW graph layers. Query enters at the top (coarsest), greedily descends to the nearest node, then refines at each lower layer. Search cost is O(log n) instead of O(n).\nDedicated vector stores built around HNSW handle graph construction and search on separate threads, persist the graph independently from the payload store, and support filtered search without degrading ANN recall — how many of the true nearest neighbors the approximate search actually returns (filters discard candidates before graph traversal completes, shrinking recall).\ngraph LR subgraph \"ANN recall example\" T[\"True top-5 neighbors\\n[A, B, C, D, E]\"] R[\"Returned by HNSW\\n[A, B, C, D, F]\"] T --\u003e|\"recall = 4/5 = 80%\"| R end graph LR subgraph \"ANN recall example\" T[\"True top-5 neighbors\\n[A, B, C, D, E]\"] R[\"Returned by HNSW\\n[A, B, C, D, F]\"] T --\u003e|\"recall = 4/5 = 80%\"| R end ⤢ Recall@5 = 80% here. E was a true neighbor; F was not. Filtered search (e.g. kind=function only) shrinks the candidate set, making misses more likely.\nSQLite-vec serialises everything through SQLite\u0026rsquo;s WAL (Write-Ahead Log) — providing crash safety but queuing concurrent writers behind each other.\nsequenceDiagram participant W1 as Writer 1 participant W2 as Writer 2 participant WAL as WAL file participant DB as Main DB file W1-\u003e\u003eWAL: append write W2-\u003e\u003eWAL: wait (locked) WAL-\u003e\u003eDB: checkpoint flush W2-\u003e\u003eWAL: append write (now unblocked) sequenceDiagram participant W1 as Writer 1 participant W2 as Writer 2 participant WAL as WAL file participant DB as Main DB file W1-\u003e\u003eWAL: append write W2-\u003e\u003eWAL: wait (locked) WAL-\u003e\u003eDB: checkpoint flush W2-\u003e\u003eWAL: append write (now unblocked) ⤢ SQLite WAL serialises concurrent writers. Fine for single-process use; degrades when indexer and query server write simultaneously at scale.\nAt small scale the difference is invisible. Across a large monorepo (thousands of files, millions of chunks), query latency climbs.\nThe macOS MCP server is currently broken.\nIssue #136 and Issue #152 document the same failure: after installing, the MCP server does not connect on macOS. The root cause is a POSIX shebang detection problem — posix_spawn requires bytes 0–1 to be #!, but a previous PR removed the shebang from scripts/run.cmd to fix Windows cmd.exe stdout cleanliness. The two constraints are mutually exclusive in a polyglot file.\nPR #154 (open at time of writing) fixes this by routing Claude Code and Cursor through a Node.js dispatcher (run.cjs) that invokes run.sh on Unix and run.bat on Windows. Until that merges, the workaround from issue #152 applies: edit plugin.json to point command at scripts/run.sh instead of scripts/run.cmd.\nLumen is the right direction. It is not stable enough yet for teams that cannot tolerate manual workarounds on install.\nruvnet/ruflo Ruflo takes the opposite approach. Where Lumen is a focused indexing tool, Ruflo is an entire orchestration platform: 100+ specialized agents, multi-agent swarms, federation with mTLS, a built-in vector database (AgentDB with HNSW), 27 hooks, cost tracking, PII detection, and support for five LLM providers.\nOn paper, it solves everything at once. In practice, that is the problem.\nRuflo suffers from the same pattern this article argues against: instead of composing small, well-understood tools, it builds a new layer on top of everything. The HNSW-backed AgentDB is faster than brute-force search, but you are running an entire orchestration platform to get what Lumen gives you with a single binary and a plugin install.\nThe token usage story is telling. Ruflo claims \u0026ldquo;75% API cost reduction\u0026rdquo; through intelligent routing and agent specialization. But coordinating swarms of agents — planning, inter-agent communication, orchestration overhead — consumes tokens too. Trading exploration tax for coordination tax is not obviously a win. It is tokenmaxxing with extra steps.\nFor a small codebase or a single focused project, Ruflo is solving problems you do not have. For a large enterprise deployment with genuine multi-agent needs, the complexity may be justified — but at that scale you have the engineering capacity to manage it.\nThe signal worth watching: Ruflo benchmarks well on SWE-Bench (84.8%). The approach is not wrong. It is aimed at a different problem than the one most developers face when their agent spends 49% of its time searching.\nRunning Local Embeddings with mlx-lm I\u0026rsquo;m currently experimenting with a MacBook Pro M5 pro with 48GB unified memory, you do not need to call an external API for embeddings. You can run better models locally than what most hosted RAG stacks use by default.\nmlx-openai-server is the right tool here. Standard GGUF runners (Ollama, llama.cpp) are 3–5× slower on M-series chips than MLX-native execution, which maps directly to the Neural Engine and unified memory. This specific server provides an Embeddings endpoint in an OpenAI API compatible structure. I\u0026rsquo;d use mlx-lm directly, but this actually worked.\nuv tool install mlx-openai-server # Serve an embedding-capable model mlx-openai-server launch \\ --model-type embeddings \\ --model-path mlx-community/Qwen3-Embedding-8B-4bit-DWQ Then point your indexer at http://localhost:8080/v1/embeddings using the OpenAI-compatible endpoint. Qwen3-8B in 4-bit DWQ fits in ~6GB of unified memory and produces 4096-dimension embeddings — substantially richer than nomic\u0026rsquo;s 768 dimensions.\nThe catch: many tools misidentify hybrid decoder models as chat-only. LM Studio, for instance, will not expose the embeddings endpoint for models it classifies as \u0026ldquo;LLM\u0026rdquo; rather than \u0026ldquo;Embedding Model.\u0026rdquo; Had this issue with Qwen3-embeddings.\nWhat Needs to Be Better The tooling exists. The patterns are proven. But the current state of the ecosystem has sharp edges:\nModel classification is broken. Hybrid models that function as both chat and embedding models get misclassified by most GUI tools. There is no standard signal for \u0026ldquo;this model supports /v1/embeddings.\u0026rdquo; You discover it by trying and failing.\nChunking is still mostly naive. Most open-source indexers use sliding character windows. Lumen\u0026rsquo;s AST-aware approach is the exception. Function-level chunks that respect language syntax are meaningfully better for code retrieval — a grep that returns the full function body outperforms one that returns lines 47–89 of an arbitrary split.\nTree shaking is a related but distinct idea from frontend build tooling — it eliminates dead code at bundle time by statically analyzing the import graph. Same goal (less junk in the output), different target: the compiled bundle, not the retrieval index. Worth naming because the terms get conflated when people discuss \u0026ldquo;pruning\u0026rdquo; what the agent sees.\ngraph LR subgraph \"Before tree shaking\" E[entry.js] --\u003e A[utils/format.js] E --\u003e B[utils/parse.js] E --\u003e C[utils/deadcode.js] A --\u003e D[lib/core.js] B --\u003e D C --\u003e D end subgraph \"After tree shaking\" E2[entry.js] --\u003e A2[utils/format.js] E2 --\u003e B2[utils/parse.js] A2 --\u003e D2[lib/core.js] B2 --\u003e D2 end graph LR subgraph \"Before tree shaking\" E[entry.js] --\u003e A[utils/format.js] E --\u003e B[utils/parse.js] E --\u003e C[utils/deadcode.js] A --\u003e D[lib/core.js] B --\u003e D C --\u003e D end subgraph \"After tree shaking\" E2[entry.js] --\u003e A2[utils/format.js] E2 --\u003e B2[utils/parse.js] A2 --\u003e D2[lib/core.js] B2 --\u003e D2 end ⤢ Tree shaking removes utils/deadcode.js — never called from the entry point. The bundler traces imports statically and drops unreachable modules. This is compile-time pruning, not retrieval-time pruning.\nNo standard MCP schema for code search. Every indexer invents its own tool names and return shapes. search_code, find_symbol, semantic_search — same operation, different contracts. The LSP standardization moment for AI tools has not happened yet.\nIndex drift under heavy editing. PostToolUse hooks sync one file at a time. A large refactor touching 40 files leaves the index inconsistent until the next SessionStart re-index. Background file watchers (like Lumen\u0026rsquo;s daemon) solve this but add process complexity.\nCost of the first index. On a large monorepo, initial indexing is slow and can be expensive if using a hosted embedding API. Local MLX models eliminate the API cost but the wall-clock time remains. There is no good incremental-from-scratch story yet.\nOr Maybe the Whole Approach Is Wrong Before concluding that the harness is the answer, it is worth holding the contrarian position: subQ argues the harness is a workaround for an architectural defect, not a solution.\nTheir thesis: RAG pipelines, chunking strategies, and retrieval systems exist because transformer attention is quadratic — every token compared against every token, cost exploding as context grows. In a sequence of N tokens, standard attention computes N² comparisons per layer. Double the context, quadruple the cost. At 100K tokens that\u0026rsquo;s 10 billion operations per layer.\nxychart-beta title \"Attention compute vs context length\" x-axis [\"1K\", \"4K\", \"16K\", \"32K\", \"64K\", \"100K\"] y-axis \"Relative compute (×)\" 0 --\u003e 10000 line [1, 16, 256, 1024, 4096, 10000] xychart-beta title \"Attention compute vs context length\" x-axis [\"1K\", \"4K\", \"16K\", \"32K\", \"64K\", \"100K\"] y-axis \"Relative compute (×)\" 0 --\u003e 10000 line [1, 16, 256, 1024, 4096, 10000] ⤢ Quadratic scaling. 1K tokens = 1× baseline. 100K tokens = 10,000× baseline. This is why RAG exists — it keeps the effective context small.\nThe industry adapted by building retrieval around the model rather than fixing the model. As they put it: \u0026ldquo;Developers and investors spend more of their time and money on workarounds than on the problem itself.\u0026rdquo;\nThe workaround stack. Source: subQ — Introducing subQ\nSubQ\u0026rsquo;s answer is a subquadratic architecture that reduces attention compute by ~1,000x at 12M tokens — making it theoretically possible to drop the entire codebase into context rather than retrieving fragments of it.\nIf that architecture holds at production scale, the harness layer becomes thinner: less indexing infrastructure, less retrieval tuning, less chunking strategy. The exploration tax disappears not because you built a better index but because the model no longer needs one.\nThat future may come. It is not here yet. Until inference cost and context reliability at multi-million-token scale match what a well-tuned retrieval harness delivers today, the harness is the practical answer. But subQ\u0026rsquo;s framing is useful: every layer of your harness is a bet that the model will not outgrow the need for it. Some of those bets will lose.\nWe\u0026rsquo;ll see.\nFor the Harness reality now, the Pattern that matters The harness is not the model. The model is a commodity. What separates a working AI coding setup from an expensive one is the infrastructure around it — the index that gives it situational awareness, the hooks that keep that index current, the sensors that catch mistakes before they land in production.\nLess tokens. More signal. That is the job.\nEngineers who understand what is underneath their tools do not reach for more context — they reach for better context. The harness is how you stop paying the exploration tax and start getting work done.\nThe tools are there, the \u0026ldquo;glue\u0026rdquo; is still largely missing.\n","permalink":"https://syndbg.dev/posts/2026-05-08-coding-harness-less-is-more-we-still-reinvent-the-wheel/","summary":"\u003ch2 id=\"what-is-a-coding-harness\"\u003eWhat Is a Coding Harness?\u003c/h2\u003e\n\u003cp\u003e\u003cstrong\u003eAgent = Model + Harness\u003c/strong\u003e. The model is the reasoning engine. The harness is everything else — the guides that steer it before it acts, the sensors that catch it after it does. (\u003ca href=\"https://martinfowler.com/articles/harness-engineering.html\"\u003eHarness Engineering\u003c/a\u003e).\u003c/p\u003e\n\u003cp\u003eIn practice a coding harness has several moving parts:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eContext injection\u003c/strong\u003e — what the agent knows before it writes a line. CLAUDE.md files, system prompts, project conventions.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eCodebase indexing\u003c/strong\u003e — a queryable map of the repo. Symbols, call graphs, semantic embeddings.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eTool layer\u003c/strong\u003e — what the agent can \u003cem\u003edo\u003c/em\u003e. Read files, run tests, call linters, search the index.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eFeedback loop\u003c/strong\u003e — sensors that observe output quality. Type errors, failing tests, lint violations, architecture drift.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eMemory / sync\u003c/strong\u003e — persistence across sessions. Hooks that re-index on session start, sync on file write.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cdiv class=\"mermaid-wrap\"\u003e\n  \u003ctemplate class=\"mermaid-source\"\u003e\u003cpre\u003egraph TD\n    A[Developer Intent] --\u003e B[Context Injection\u003cbr/\u003eCLAUDE.md · system prompt · conventions]\n    B --\u003e C[Model]\n    C --\u003e D[Tool Layer\u003cbr/\u003eread · write · run tests · search index]\n    D --\u003e E[Codebase Index\u003cbr/\u003esymbols · embeddings · call graph]\n    E --\u003e|semantic search results| C\n    D --\u003e F[Feedback Loop\u003cbr/\u003etype errors · lint · test failures]\n    F --\u003e|sensor output| C\n    C --\u003e G[Output]\n    G --\u003e H[Memory / Sync\u003cbr/\u003ePostToolUse hook · incremental re-index]\n    H --\u003e E\u003c/pre\u003e\u003c/template\u003e\n  \u003cpre class=\"mermaid-fallback\"\u003egraph TD\n    A[Developer Intent] --\u003e B[Context Injection\u003cbr/\u003eCLAUDE.md · system prompt · conventions]\n    B --\u003e C[Model]\n    C --\u003e D[Tool Layer\u003cbr/\u003eread · write · run tests · search index]\n    D --\u003e E[Codebase Index\u003cbr/\u003esymbols · embeddings · call graph]\n    E --\u003e|semantic search results| C\n    D --\u003e F[Feedback Loop\u003cbr/\u003etype errors · lint · test failures]\n    F --\u003e|sensor output| C\n    C --\u003e G[Output]\n    G --\u003e H[Memory / Sync\u003cbr/\u003ePostToolUse hook · incremental re-index]\n    H --\u003e E\u003c/pre\u003e\n  \u003cdiv class=\"mermaid-frame-host\"\u003e\u003c/div\u003e\n  \u003cbutton type=\"button\" class=\"mermaid-expand-btn\" aria-label=\"Expand diagram\"\u003e⤢\u003c/button\u003e\n\u003c/div\u003e\n\n\u003cp\u003eThe vocabulary here is not new. Russell \u0026amp; Norvig’s \u003cem\u003eAI: A Modern Approach\u003c/em\u003e (1995) defined an agent as anything that perceives its environment through \u003cstrong\u003esensors\u003c/strong\u003e and acts upon it through \u003cstrong\u003eactuators\u003c/strong\u003e. Chip Huyen applies this directly to modern LLM agents in \u003ca href=\"https://huyenchip.com/2025/01/07/agents.html\"\u003eAgents\u003c/a\u003e (2025): the model is the brain, tools are the actuators, observations are the sensors. Böckeler maps the same structure onto coding harnesses — guides are actuators (steer behavior forward), sensors are feedback mechanisms (observe what came out).\u003c/p\u003e","title":"Coding Harness, Less Is More, We Still Re-invent the Wheel inefficiently"},{"content":"It\u0026rsquo;s 2026, and everyone is claiming to run parallel agents through one harness or another: Claude Code, Cursor, Aider, and whatnot.\nStill, I think we need to understand how these tools work underneath, and even use the same primitives ourselves in the old and boring way: manually.\nThat was my experience with Git worktrees. I knew about them for more than 10 years, but only started using them seriously last year.\nThe goal of this post is to share my experience with Git worktrees from a developer perspective: how they help in normal day-to-day work, how AI agents use them, what CLI tools are available, and where it makes sense to use git worktree instead of something like git bisect, which is what I used to reach for first. You can read the official Linux man page-style documentation too, but this post is usage-centered, not a CLI reference.\nThat changes when local state starts to matter. You have a server running, generated files, uncommitted debugging changes, a database migration half-tested, and three editor tabs open in the middle of a problem. Then someone asks you to review a PR, or a production bug appears, or you want to run an AI coding agent on a separate task.\nThis is the point where branch switching stops being cheap.\nWorktrees solve the problem at the right layer: one repository, multiple working directories, each on its own branch. No duplicate clone, no shared working directory, no stash dance.\nWhat is a Git Worktree? A Git worktree lets you check out more than one branch from the same repository at the same time. Each branch gets its own directory. All of them share the same underlying Git object database.\nThe mental model is simple:\nrepo/ main branch repo-review-reconciler/ review/reconciler-backoff branch repo-hotfix-worker/ fix/worker-timeout branch Instead of changing the branch inside one directory, you change directories.\nGit added worktrees in Git 2.5, released in 2015. This is not a new experimental feature. It has been sitting there quietly for almost a decade, mostly used by people who got tired of stashing at exactly the wrong time or running git bisect to oblivion to catch issues.\nThe important rule:\nOne branch can only be checked out in one worktree at a time.\nGit will not let you check out the same branch in two directories simultaneously. That is a feature, not a limitation. It prevents two working directories from trying to mutate the same branch state.\nArchitecture Overview A normal clone gives you one working directory attached to one checked-out branch. A repository with worktrees gives you many working directories attached to the same Git history.\ngraph TB subgraph GIT[\"Shared Git Repository\"] OBJ[\".git object database\"] REM[\"origin/main, origin/feature, origin/fix\"] end subgraph WT1[\"app/\"] MAIN[\"branch: main\"] DEV[\"local server running\"] end subgraph WT2[\"service-review-checkout/\"] REVIEW[\"branch: review/reconciler-backoff\"] TESTS[\"run tests for PR\"] end subgraph WT3[\"service-hotfix-worker/\"] HOTFIX[\"branch: fix/worker-timeout\"] DEBUG[\"debug production issue\"] end WT1 --\u003e OBJ WT2 --\u003e OBJ WT3 --\u003e OBJ OBJ --\u003e REM graph TB subgraph GIT[\"Shared Git Repository\"] OBJ[\".git object database\"] REM[\"origin/main, origin/feature, origin/fix\"] end subgraph WT1[\"app/\"] MAIN[\"branch: main\"] DEV[\"local server running\"] end subgraph WT2[\"service-review-checkout/\"] REVIEW[\"branch: review/reconciler-backoff\"] TESTS[\"run tests for PR\"] end subgraph WT3[\"service-hotfix-worker/\"] HOTFIX[\"branch: fix/worker-timeout\"] DEBUG[\"debug production issue\"] end WT1 --\u003e OBJ WT2 --\u003e OBJ WT3 --\u003e OBJ OBJ --\u003e REM ⤢ The directories are independent at the filesystem level. They can have different build artifacts, different uncommitted changes, different running processes, and different editor windows.\nThe history is still shared. Fetch once, push normally, merge normally.\nBasic Commands Native Git has everything you need. The full command reference is in the git worktree documentation:\n# Create a worktree for an existing branch git worktree add ../service-reconciler feature/reconciler-backoff # Create a new branch and a worktree for it git worktree add -b fix/worker-timeout ../service-fix-worker-timeout main # Create a worktree from a remote branch git fetch origin git worktree add ../service-review-migrations origin/review/migration-locking # List active worktrees git worktree list # Remove a worktree after you are done git worktree remove ../service-review-migrations # Clean up stale metadata if a directory was removed manually git worktree prune There is no magic beyond that. Each worktree is just another directory. You can cd into it, open it in your editor, run tests, start a dev server, or hand it to an agent.\nThe related Git commands worth knowing are git fetch, git branch, and git status. Worktrees do not replace those commands; they give each branch a separate working directory.\nWhy Worktrees Matter for Developers The classic worktree example is the urgent hotfix. You are halfway through a feature and production breaks. Without worktrees, you either stash your changes, commit something unfinished, or clone the repository again somewhere else.\nWith worktrees:\ngit worktree add -b fix/worker-timeout ../service-fix-worker-timeout main cd ../service-fix-worker-timeout Your feature branch stays exactly where it was. Your editor, running server, generated files, and local debugging state do not move. The hotfix gets a clean directory from main.\nThat is useful even if you never touch AI tools.\nReviewing a PR locally Some reviews can be done from a diff. Some cannot. If the change touches build tooling, database migrations, generated clients, controller logic, or anything with runtime behavior, you usually need to run it.\ngit fetch origin git worktree add ../service-review-migration-locking origin/pr/migration-locking cd ../service-review-migration-locking go test ./... make integration-test go run ./cmd/api --listen :8081 Your own branch keeps running in the original directory. The review has a clean checkout and its own dependency state. When the review is done:\ncd ../service git worktree remove ../service-review-migration-locking No branch switching. No cleanup from someone else\u0026rsquo;s generated files. No accidental commit of review-only changes into your feature branch.\nComparing two branches side by side Sometimes the diff is not enough. You want to run the old behavior and the new behavior at the same time.\ngit worktree add ../service-before-release release/2026-04 git worktree add ../service-after-release main Now you can start both versions on different ports:\n# service-before-release go run ./cmd/api --listen :8080 # service-after-release go run ./cmd/api --listen :8081 This is especially useful for API behavior changes, migration behavior, queue processing, controller reconciliation, and performance comparisons. You can keep both branches running, send the same request or replay the same message, and see what actually changed.\nDebugging two related repositories Worktrees become even more useful when the bug crosses repository boundaries. A common backend example is a service and the Go client, protobuf definitions, or shared platform library it depends on.\nservice/ main branch service-debug-search-timeout/ fix/search-timeout branch client-go/ main branch client-go-debug-search-timeout/ fix/search-timeout-context branch You can run the service fix in one worktree and the client or shared library fix in another:\n# Service repository cd ~/src/service git worktree add -b fix/search-timeout ../service-debug-search-timeout main # Go client repository cd ~/src/client-go git worktree add -b fix/search-timeout-context ../client-go-debug-search-timeout main Now the debugging environment is explicit:\nservice-debug-search-timeout -\u0026gt; service on :8080 client-go-debug-search-timeout -\u0026gt; client tests against :8080 service and client-go main dirs -\u0026gt; untouched If the investigation goes nowhere, delete both worktrees. If it works, keep the branches and open two PRs. Your normal development setup remains intact.\nGit Bisect vs Worktrees git bisect and git worktree solve different debugging problems.\ngit bisect is for finding the commit that introduced a regression. You give Git one known-good commit and one known-bad commit, then Git checks out commits in the middle until the bad commit is isolated.\ngit bisect start git bisect bad HEAD git bisect good v1.42.0 # Run your test at each step go test ./internal/search -run TestSearchTimeout git bisect good # or git bisect bad git bisect reset Use git bisect when the question is:\nWhich commit broke this?\nUse worktrees when the question is:\nHow do I investigate this without destroying my current working state?\nThe two tools also work well together. If you are in the middle of feature work and need to bisect a production regression, create a clean worktree and run the bisect there:\ngit worktree add ../service-bisect main cd ../service-bisect git bisect start git bisect bad HEAD git bisect good v1.42.0 That keeps the bisect checkout churn away from your main development directory. This matters because git bisect repeatedly checks out different commits. If your service needs generated protobuf clients, local config, migrations, test containers, or a running dependency, doing that in your main checkout is painful.\nA simple rule:\nUse worktrees when you need multiple working directories at the same time. Use bisect when you need Git to search history for the first bad commit. Use bisect inside a worktree when you need to debug history without interrupting current work. For quick regressions with a reliable test command, git bisect is usually easier. For messy investigations where you need to compare branches, keep servers running, inspect a PR, or preserve local state, worktrees are easier. For serious production debugging, I usually start with a worktree so the investigation has its own clean room, then use bisect inside it if I need the exact breaking commit.\nWhen to Use Worktrees Use a worktree when the cost of switching branches is higher than the cost of creating another directory.\nGood use cases:\nYou are in the middle of feature work and need to handle an urgent bug. You need to review a PR locally without disturbing your current branch. You want to compare two branches by running both at the same time. You are testing a dependency upgrade and expect generated files or lockfile churn. You are debugging behavior across two services or two repositories. You want to run git bisect without disturbing your current checkout. You want to give an AI coding agent its own isolated workspace. You need a clean release branch checkout while continuing work on main. You probably do not need a worktree for a two-minute edit, a simple documentation change, or a branch you will check out once and immediately merge. Normal branch switching is still fine when local state does not matter.\nThe signal is friction. If you are about to stash, stop a server, save a half-written note, or worry about generated files, a worktree is probably cleaner.\nWhy Worktrees Matter for AI Agents An AI coding agent is just another writer in your working directory. If you run one agent in your main checkout, everything is fine. One branch, one filesystem, one stream of changes.\nThe moment you run two agents against the same directory, the model breaks.\nThey can edit the same file. They can race on go.mod, generated protobuf code, migrations, or OpenAPI clients. One agent can run formatting while another is reading stale state. They can overwrite each other\u0026rsquo;s generated files. Even if the agents are careful, the working directory is not designed for concurrent mutation.\nWorktrees give each agent a separate branch and a separate filesystem:\nsequenceDiagram participant Dev as Developer participant WT1 as worktree: feature/reconciler-backoff participant A1 as Agent 1 participant WT2 as worktree: fix/integration-tests participant A2 as Agent 2 participant Git as Shared Git History Dev-\u003e\u003eWT1: create feature/reconciler-backoff worktree Dev-\u003e\u003eWT2: create fix/integration-tests worktree Dev-\u003e\u003eA1: improve controller retry behavior Dev-\u003e\u003eA2: fix flaky integration tests A1-\u003e\u003eWT1: edit files, run tests A2-\u003e\u003eWT2: edit files, run tests WT1-\u003e\u003eGit: commit branch changes WT2-\u003e\u003eGit: commit branch changes Dev-\u003e\u003eGit: review and merge intentionally sequenceDiagram participant Dev as Developer participant WT1 as worktree: feature/reconciler-backoff participant A1 as Agent 1 participant WT2 as worktree: fix/integration-tests participant A2 as Agent 2 participant Git as Shared Git History Dev-\u003e\u003eWT1: create feature/reconciler-backoff worktree Dev-\u003e\u003eWT2: create fix/integration-tests worktree Dev-\u003e\u003eA1: improve controller retry behavior Dev-\u003e\u003eA2: fix flaky integration tests A1-\u003e\u003eWT1: edit files, run tests A2-\u003e\u003eWT2: edit files, run tests WT1-\u003e\u003eGit: commit branch changes WT2-\u003e\u003eGit: commit branch changes Dev-\u003e\u003eGit: review and merge intentionally ⤢ That gives you the right failure boundary. If one agent goes in the wrong direction, delete that worktree. If another produces something useful, review the branch and merge it.\nThe review step is still human. The isolation lets you parallelize the execution without pretending that all generated code is automatically correct.\nThe Agent Workflow The simplest version uses native Git:\ngit worktree add -b agent/reconciler-backoff ../service-agent-reconciler main git worktree add -b agent/storage-timeout ../service-agent-storage main git worktree add -b agent/fix-integration-tests ../service-agent-tests main Then run one agent per directory:\ncd ../service-agent-reconciler claude \u0026#34;Add exponential backoff to the controller reconcile loop and cover it with tests\u0026#34; cd ../service-agent-storage cursor-agent \u0026#34;Investigate object storage timeout handling and make retries idempotent\u0026#34; cd ../service-agent-tests codex \u0026#34;Fix the flaky integration tests around postgres advisory locks\u0026#34; The exact tool does not matter. Cursor, Claude Code, Codex, or any local agent benefits from the same isolation model.\nAfter the agents finish, review like any other branch:\ncd ../service-agent-reconciler git status git diff main...HEAD go test ./... Then merge, cherry-pick, or throw it away.\nUsing Worktrunk Native commands work, but the lifecycle has friction. You create a branch, create a directory, switch into it, run setup commands, maybe copy .env, maybe install dependencies, then later merge and clean up.\nWorktrunk (wt) wraps that flow:\nbrew install max-sixty/tap/worktrunk # or cargo install worktrunk Add shell integration so wt switch can actually change your directory:\n# .zshrc eval \u0026#34;$(wt shell-init zsh)\u0026#34; The core workflow becomes:\n# Create a new branch and worktree, then cd into it wt switch --create feature-reconciler-backoff # See worktrees with status wt list # Merge back to main wt merge main # Remove the worktree and branch wt remove Worktrunk can also launch a command directly inside the new worktree:\nwt switch -x claude -c feature-reconciler-backoff -- \u0026#34;Add exponential backoff to the controller reconcile loop\u0026#34; The useful part is not just shorter syntax. It is lifecycle management. A good worktree workflow needs hooks for local setup:\n# .wt.toml [hooks] post_create = [ \u0026#34;cp ../service/.env .env\u0026#34;, \u0026#34;go mod download\u0026#34;, \u0026#34;make generate\u0026#34; ] That is where worktrees become comfortable. New branch, clean directory, local setup done, agent or developer starts working.\nTooling Around Worktrees There is a small but growing ecosystem around this because the primitive is simple and the workflows around it are repetitive.\nTool Type Best for Worktrunk (wt) CLI Lifecycle management, hooks, agent launching, scripting git-worktree-runner (gtr) CLI Editor integration, CodeRabbit workflows, AI tool support Claude Squad tmux wrapper Keyboard-first users who want multiple agents in tmux Daintree Desktop app Visibility into agent state across worktrees Superset IDE Full multi-agent orchestration with a dashboard gwq CLI Lightweight status dashboard with tmux integration My default is still a terminal-first workflow, complimented with more and more usage of wt. So, to understand what these CLI tools abstract away, start with native git worktree for a few days. If you find yourself repeating setup and cleanup steps, add wt. If you are running enough agents that you no longer know which one is waiting for input, add a visibility layer.\nCursor and Claude Code also have first-class worktree now. That is the important shift: the old Git feature became a normal part of the AI coding toolchain.\nCommon Pitfalls Trying to check out the same branch twice. Git prevents this. Create a new branch per worktree, even if it is temporary.\ngit worktree add -b review/migration-locking ../service-review-migration-locking origin/migration-locking Deleting the directory manually. If you remove a worktree with rm -rf, Git may still keep metadata for it. Prefer:\ngit worktree remove ../service-review-migration-locking If the directory is already gone:\ngit worktree prune Forgetting local setup. A worktree does not automatically have your copied .env, downloaded modules, generated clients, local database state, or running services. Automate that setup if you create worktrees often.\nSharing ports between worktrees. Two copies of the same service cannot both bind to localhost:8080. Run them on different ports when comparing branches.\nLetting worktrees pile up. Worktrees are cheap, but not free. List them regularly:\ngit worktree list Remove the branches you no longer need.\nTreating agent output as merged work. A worktree is isolation, not validation. You still need to read the diff, run tests, check generated files, and decide whether the change belongs in the main branch.\nWhat You End Up With After adopting worktrees, what you have is:\nA stable main development directory that does not get interrupted by every review or hotfix Clean checkouts for PR reviews and release branches A practical way to compare two versions of the same service side by side Isolated directories for debugging multi-repo problems A safe filesystem boundary for running multiple AI agents in parallel The feature itself is small. The workflow change is not. You stop treating the working directory as the only place work can happen.\nFor developers, that means fewer stashes, cleaner reviews, and less context loss. For AI agents, it means parallel work without pretending one mutable directory can safely host multiple writers.\nStart with git worktree add. Use it for your next local PR review or urgent bugfix. Once the pattern sticks, add tooling around the lifecycle.\n","permalink":"https://syndbg.dev/posts/2026-05-02-git-worktrees-parallel-development-without-cloning-everything-twice/","summary":"\u003cp\u003eIt\u0026rsquo;s 2026, and everyone is claiming to run parallel agents through one harness or another: Claude Code, Cursor, Aider, and whatnot.\u003c/p\u003e\n\u003cp\u003eStill, I think we need to understand how these tools work underneath, and even use the same primitives ourselves in the old and boring way: manually.\u003c/p\u003e\n\u003cp\u003eThat was my experience with Git worktrees. I knew about them for more than 10 years, but only started using them seriously last year.\u003c/p\u003e","title":"Git Worktrees: Parallel Development Without Cloning Everything Twice"},{"content":"Kubernetes storage is one of those areas that looks simple from the outside — you create a PersistentVolumeClaim, a pod mounts it, done. But the moment you need to integrate your own storage backend, you\u0026rsquo;re staring at the Container Storage Interface spec, sidecar containers you\u0026rsquo;ve never heard of, and gRPC services that have to be wired together just right.\nI had the pleasure of writing and contributing to a production-grade CSI driver end-to-end. This post covers everything I wish I had in one place: what CSI actually is, how Kubernetes orchestrates it, and how to implement all three services in Go.\nWhat is CSI? CSI (Container Storage Interface) is a standardized gRPC API between Kubernetes and storage providers. Before CSI existed, storage drivers were compiled directly into Kubernetes — adding a new one meant patching the Kubernetes tree. If you remember a few Kubernetes versions ago where your non-Azure Kubernetes is actually logging non-stop that Azure is not working? Yeah, that kind of problem due to coupling.\nCSI moved drivers out-of-tree: your storage backend ships its own binary, Kubernetes talks to it over a Unix socket.\nThe spec defines three gRPC services:\nIdentityService — who are you, are you healthy? ControllerService — manage volumes at the infrastructure level (create, delete, attach, snapshot) NodeService — manage volumes on a specific node (format, mount, bind-mount into pods) Your driver implements these. Kubernetes provides sidecar containers that translate Kubernetes events into the right RPC calls.\nArchitecture Overview Two separate binaries run in the cluster:\ngraph TB subgraph CP[\"Controller Pod — Deployment, 1 replica\"] PROV[csi-provisioner] ATT[csi-attacher] SNAP[csi-snapshotter] RES[csi-resizer] CLIVE[liveness-probe] CDRV[\"Your Driver\\nControllerService + IdentityService\"] PROV --\u003e|gRPC unix socket| CDRV ATT --\u003e|gRPC unix socket| CDRV SNAP --\u003e|gRPC unix socket| CDRV RES --\u003e|gRPC unix socket| CDRV CLIVE --\u003e|gRPC unix socket| CDRV end subgraph NP[\"Node Pod — DaemonSet, 1 per node\"] REG[node-driver-registrar] NLIVE[liveness-probe] NDRV[\"Your Driver\\nNodeService + IdentityService\"] REG --\u003e|gRPC unix socket| NDRV NLIVE --\u003e|gRPC unix socket| NDRV end KBL[kubelet] --\u003e|gRPC unix socket| NDRV REG --\u003e|registers socket path| KBL graph TB subgraph CP[\"Controller Pod — Deployment, 1 replica\"] PROV[csi-provisioner] ATT[csi-attacher] SNAP[csi-snapshotter] RES[csi-resizer] CLIVE[liveness-probe] CDRV[\"Your Driver\\nControllerService + IdentityService\"] PROV --\u003e|gRPC unix socket| CDRV ATT --\u003e|gRPC unix socket| CDRV SNAP --\u003e|gRPC unix socket| CDRV RES --\u003e|gRPC unix socket| CDRV CLIVE --\u003e|gRPC unix socket| CDRV end subgraph NP[\"Node Pod — DaemonSet, 1 per node\"] REG[node-driver-registrar] NLIVE[liveness-probe] NDRV[\"Your Driver\\nNodeService + IdentityService\"] REG --\u003e|gRPC unix socket| NDRV NLIVE --\u003e|gRPC unix socket| NDRV end KBL[kubelet] --\u003e|gRPC unix socket| NDRV REG --\u003e|registers socket path| KBL ⤢ The controller pod runs centrally and manages volume lifecycle against your storage backend API. The node pod runs on every machine and handles the actual OS-level work: formatting disks, mounting filesystems, bind-mounting into pods.\nThey communicate via Unix domain sockets, not TCP. Each sidecar and the driver share an emptyDir (controller) or hostPath (node) volume where the socket lives.\nHow Kubernetes Orchestrates CSI The sidecars are the glue. You don\u0026rsquo;t call your driver directly — the sidecars watch Kubernetes API resources and translate events into gRPC calls.\nSidecar Watches Triggers RPC csi-provisioner PersistentVolumeClaim CreateVolume / DeleteVolume csi-attacher VolumeAttachment ControllerPublishVolume / ControllerUnpublishVolume csi-snapshotter VolumeSnapshot CreateSnapshot / DeleteSnapshot csi-resizer PVC resize ControllerExpandVolume node-driver-registrar startup registers socket with kubelet kubelet Pod scheduling NodeStageVolume / NodePublishVolume kubelet talks to the node plugin directly after node-driver-registrar registers the socket path at /var/lib/kubelet/plugins_registry/.\nThe Full Volume Lifecycle Provision → Mount sequenceDiagram actor User participant K8s as Kubernetes API participant PROV as csi-provisioner participant ADC as AttachDetachController participant ATT as csi-attacher participant CTRL as ControllerService participant BACK as Storage Backend participant KBL as kubelet participant NODE as NodeService User-\u003e\u003eK8s: create PVC (16Gi, RWO) K8s-\u003e\u003ePROV: PVC event PROV-\u003e\u003eCTRL: CreateVolume(name, 16Gi, topology) CTRL-\u003e\u003eBACK: create volume CTRL-\u003e\u003eBACK: poll status → \"available\" CTRL--\u003e\u003ePROV: CreateVolumeResponse{volumeID, capacityBytes, topology} PROV-\u003e\u003eK8s: create PersistentVolume{volumeHandle=volumeID} K8s-\u003e\u003eK8s: bind PV ↔ PVC User-\u003e\u003eK8s: create Pod using PVC K8s-\u003e\u003eADC: Pod scheduled to node ADC-\u003e\u003eK8s: create VolumeAttachment{volumeID, nodeID} K8s-\u003e\u003eATT: VolumeAttachment event ATT-\u003e\u003eCTRL: ControllerPublishVolume(volumeID, nodeID) CTRL-\u003e\u003eBACK: attach volume to VM BACK--\u003e\u003eCTRL: attached, devicePath=/dev/vdb CTRL--\u003e\u003eATT: ControllerPublishVolumeResponse{PublishContext{DevicePath=/dev/vdb}} ATT-\u003e\u003eK8s: VolumeAttachment.status.attached=true KBL-\u003e\u003eNODE: NodeStageVolume(volumeID, stagingPath, PublishContext) NODE-\u003e\u003eNODE: mkfs.ext4 /dev/vdb NODE-\u003e\u003eNODE: mount /dev/vdb → stagingPath NODE--\u003e\u003eKBL: NodeStageVolumeResponse{} KBL-\u003e\u003eNODE: NodePublishVolume(volumeID, targetPath, stagingPath) NODE-\u003e\u003eNODE: mount --bind stagingPath → targetPath NODE--\u003e\u003eKBL: NodePublishVolumeResponse{} K8s--\u003e\u003eUser: Pod Running, /data mounted sequenceDiagram actor User participant K8s as Kubernetes API participant PROV as csi-provisioner participant ADC as AttachDetachController participant ATT as csi-attacher participant CTRL as ControllerService participant BACK as Storage Backend participant KBL as kubelet participant NODE as NodeService User-\u003e\u003eK8s: create PVC (16Gi, RWO) K8s-\u003e\u003ePROV: PVC event PROV-\u003e\u003eCTRL: CreateVolume(name, 16Gi, topology) CTRL-\u003e\u003eBACK: create volume CTRL-\u003e\u003eBACK: poll status → \"available\" CTRL--\u003e\u003ePROV: CreateVolumeResponse{volumeID, capacityBytes, topology} PROV-\u003e\u003eK8s: create PersistentVolume{volumeHandle=volumeID} K8s-\u003e\u003eK8s: bind PV ↔ PVC User-\u003e\u003eK8s: create Pod using PVC K8s-\u003e\u003eADC: Pod scheduled to node ADC-\u003e\u003eK8s: create VolumeAttachment{volumeID, nodeID} K8s-\u003e\u003eATT: VolumeAttachment event ATT-\u003e\u003eCTRL: ControllerPublishVolume(volumeID, nodeID) CTRL-\u003e\u003eBACK: attach volume to VM BACK--\u003e\u003eCTRL: attached, devicePath=/dev/vdb CTRL--\u003e\u003eATT: ControllerPublishVolumeResponse{PublishContext{DevicePath=/dev/vdb}} ATT-\u003e\u003eK8s: VolumeAttachment.status.attached=true KBL-\u003e\u003eNODE: NodeStageVolume(volumeID, stagingPath, PublishContext) NODE-\u003e\u003eNODE: mkfs.ext4 /dev/vdb NODE-\u003e\u003eNODE: mount /dev/vdb → stagingPath NODE--\u003e\u003eKBL: NodeStageVolumeResponse{} KBL-\u003e\u003eNODE: NodePublishVolume(volumeID, targetPath, stagingPath) NODE-\u003e\u003eNODE: mount --bind stagingPath → targetPath NODE--\u003e\u003eKBL: NodePublishVolumeResponse{} K8s--\u003e\u003eUser: Pod Running, /data mounted ⤢ Three distinct paths to understand here:\nProvision (csi-provisioner → CreateVolume) — creates the volume in your backend, Kubernetes creates a PV. Attach (csi-attacher → ControllerPublishVolume) — attaches the block device to the VM running the pod. Returns the device path (e.g. /dev/vdb) in PublishContext. Stage + Publish (kubelet → NodeStageVolume + NodePublishVolume) — formats and mounts the device to a staging path, then bind-mounts that into the pod\u0026rsquo;s specific target path. The staging/publish split exists so multiple pods on the same node can share one formatted device via bind mounts, rather than formatting once per pod.\nUnmount → Delete sequenceDiagram actor User participant K8s as Kubernetes API participant PROV as csi-provisioner participant ADC as AttachDetachController participant ATT as csi-attacher participant CTRL as ControllerService participant BACK as Storage Backend participant KBL as kubelet participant NODE as NodeService User-\u003e\u003eK8s: delete Pod KBL-\u003e\u003eNODE: NodeUnpublishVolume(volumeID, targetPath) NODE-\u003e\u003eNODE: umount targetPath NODE--\u003e\u003eKBL: NodeUnpublishVolumeResponse{} KBL-\u003e\u003eNODE: NodeUnstageVolume(volumeID, stagingPath) NODE-\u003e\u003eNODE: umount stagingPath NODE--\u003e\u003eKBL: NodeUnstageVolumeResponse{} ADC-\u003e\u003eK8s: delete VolumeAttachment K8s-\u003e\u003eATT: VolumeAttachment deletion event ATT-\u003e\u003eCTRL: ControllerUnpublishVolume(volumeID, nodeID) CTRL-\u003e\u003eBACK: detach volume from VM CTRL--\u003e\u003eATT: ControllerUnpublishVolumeResponse{} User-\u003e\u003eK8s: delete PVC K8s-\u003e\u003ePROV: PVC deletion event PROV-\u003e\u003eCTRL: DeleteVolume(volumeID) CTRL-\u003e\u003eBACK: delete volume CTRL--\u003e\u003ePROV: DeleteVolumeResponse{} PROV-\u003e\u003eK8s: delete PersistentVolume sequenceDiagram actor User participant K8s as Kubernetes API participant PROV as csi-provisioner participant ADC as AttachDetachController participant ATT as csi-attacher participant CTRL as ControllerService participant BACK as Storage Backend participant KBL as kubelet participant NODE as NodeService User-\u003e\u003eK8s: delete Pod KBL-\u003e\u003eNODE: NodeUnpublishVolume(volumeID, targetPath) NODE-\u003e\u003eNODE: umount targetPath NODE--\u003e\u003eKBL: NodeUnpublishVolumeResponse{} KBL-\u003e\u003eNODE: NodeUnstageVolume(volumeID, stagingPath) NODE-\u003e\u003eNODE: umount stagingPath NODE--\u003e\u003eKBL: NodeUnstageVolumeResponse{} ADC-\u003e\u003eK8s: delete VolumeAttachment K8s-\u003e\u003eATT: VolumeAttachment deletion event ATT-\u003e\u003eCTRL: ControllerUnpublishVolume(volumeID, nodeID) CTRL-\u003e\u003eBACK: detach volume from VM CTRL--\u003e\u003eATT: ControllerUnpublishVolumeResponse{} User-\u003e\u003eK8s: delete PVC K8s-\u003e\u003ePROV: PVC deletion event PROV-\u003e\u003eCTRL: DeleteVolume(volumeID) CTRL-\u003e\u003eBACK: delete volume CTRL--\u003e\u003ePROV: DeleteVolumeResponse{} PROV-\u003e\u003eK8s: delete PersistentVolume ⤢ Exact reverse order: unpublish → unstage → detach → delete. Each step must succeed before the next begins.\nSnapshots sequenceDiagram actor User participant K8s as Kubernetes API participant SNAP as csi-snapshotter participant PROV as csi-provisioner participant CTRL as ControllerService participant BACK as Storage Backend User-\u003e\u003eK8s: create VolumeSnapshot{source=PVC} K8s-\u003e\u003eSNAP: VolumeSnapshot event SNAP-\u003e\u003eCTRL: CreateSnapshot(name, sourceVolumeID) CTRL-\u003e\u003eBACK: create snapshot CTRL-\u003e\u003eBACK: poll status → \"available\" CTRL--\u003e\u003eSNAP: CreateSnapshotResponse{snapshotID, readyToUse=true} SNAP-\u003e\u003eK8s: create VolumeSnapshotContent{snapshotHandle=snapshotID} K8s--\u003e\u003eUser: VolumeSnapshot Ready=true User-\u003e\u003eK8s: create PVC{dataSource=VolumeSnapshot} K8s-\u003e\u003ePROV: PVC event PROV-\u003e\u003eCTRL: CreateVolume(name, size, contentSource={snapshotID}) CTRL-\u003e\u003eBACK: create volume from snapshot CTRL--\u003e\u003ePROV: CreateVolumeResponse{volumeID, resizeRequired=true} Note over PROV,K8s: NodeStageVolume will resize filesystem sequenceDiagram actor User participant K8s as Kubernetes API participant SNAP as csi-snapshotter participant PROV as csi-provisioner participant CTRL as ControllerService participant BACK as Storage Backend User-\u003e\u003eK8s: create VolumeSnapshot{source=PVC} K8s-\u003e\u003eSNAP: VolumeSnapshot event SNAP-\u003e\u003eCTRL: CreateSnapshot(name, sourceVolumeID) CTRL-\u003e\u003eBACK: create snapshot CTRL-\u003e\u003eBACK: poll status → \"available\" CTRL--\u003e\u003eSNAP: CreateSnapshotResponse{snapshotID, readyToUse=true} SNAP-\u003e\u003eK8s: create VolumeSnapshotContent{snapshotHandle=snapshotID} K8s--\u003e\u003eUser: VolumeSnapshot Ready=true User-\u003e\u003eK8s: create PVC{dataSource=VolumeSnapshot} K8s-\u003e\u003ePROV: PVC event PROV-\u003e\u003eCTRL: CreateVolume(name, size, contentSource={snapshotID}) CTRL-\u003e\u003eBACK: create volume from snapshot CTRL--\u003e\u003ePROV: CreateVolumeResponse{volumeID, resizeRequired=true} Note over PROV,K8s: NodeStageVolume will resize filesystem ⤢ Volume Expansion sequenceDiagram actor User participant K8s as Kubernetes API participant RES as csi-resizer participant CTRL as ControllerService participant BACK as Storage Backend participant KBL as kubelet participant NODE as NodeService User-\u003e\u003eK8s: patch PVC 16Gi → 32Gi K8s-\u003e\u003eRES: PVC resize event RES-\u003e\u003eCTRL: ControllerExpandVolume(volumeID, 32Gi) CTRL-\u003e\u003eBACK: expand volume CTRL-\u003e\u003eBACK: poll → new size confirmed CTRL--\u003e\u003eRES: ControllerExpandVolumeResponse{nodeExpansionRequired=true} RES-\u003e\u003eK8s: update PV capacity = 32Gi KBL-\u003e\u003eNODE: NodeExpandVolume(volumeID, volumePath, 32Gi) NODE-\u003e\u003eNODE: rescan block device geometry NODE-\u003e\u003eNODE: resize2fs /dev/vdb NODE--\u003e\u003eKBL: NodeExpandVolumeResponse{} K8s--\u003e\u003eUser: PVC capacity = 32Gi sequenceDiagram actor User participant K8s as Kubernetes API participant RES as csi-resizer participant CTRL as ControllerService participant BACK as Storage Backend participant KBL as kubelet participant NODE as NodeService User-\u003e\u003eK8s: patch PVC 16Gi → 32Gi K8s-\u003e\u003eRES: PVC resize event RES-\u003e\u003eCTRL: ControllerExpandVolume(volumeID, 32Gi) CTRL-\u003e\u003eBACK: expand volume CTRL-\u003e\u003eBACK: poll → new size confirmed CTRL--\u003e\u003eRES: ControllerExpandVolumeResponse{nodeExpansionRequired=true} RES-\u003e\u003eK8s: update PV capacity = 32Gi KBL-\u003e\u003eNODE: NodeExpandVolume(volumeID, volumePath, 32Gi) NODE-\u003e\u003eNODE: rescan block device geometry NODE-\u003e\u003eNODE: resize2fs /dev/vdb NODE--\u003e\u003eKBL: NodeExpandVolumeResponse{} K8s--\u003e\u003eUser: PVC capacity = 32Gi ⤢ Implementing the Driver in Go Project Structure cmd/ controller/main.go # wires ControllerService + IdentityService node/main.go # wires NodeService + IdentityService internal/ server/server.go # Unix socket listener + gRPC server driver/ identity.go controller.go # struct + capability declarations controller_volume.go controller_snapshot.go node.go mounter.go # interface wrapping k8s.io/mount-utils metadata.go # node instance ID + AZ Two binaries, one shared internal/driver package. The controller binary registers ControllerServer + IdentityServer. The node binary registers NodeServer + IdentityServer.\nThe gRPC Server func CreateListener(endpoint string) (net.Listener, error) { path := strings.TrimPrefix(endpoint, \u0026#34;unix://\u0026#34;) _ = os.Remove(path) return net.Listen(\u0026#34;unix\u0026#34;, path) } func CreateGRPCServer(logger *slog.Logger) *grpc.Server { return grpc.NewServer( grpc.ChainUnaryInterceptor(requestLogger(logger)), ) } // cmd/controller/main.go — error handling omitted for brevity. listener, _ := internal.CreateListener(cfg.CSIEndpoint) grpcServer := internal.CreateGRPCServer(logger) csiproto.RegisterControllerServer(grpcServer, controllerService) csiproto.RegisterIdentityServer(grpcServer, identityService) grpcServer.Serve(listener) IdentityService The simplest service. Declares what the driver supports and responds to health checks.\ntype IdentityService struct { csiproto.UnimplementedIdentityServer logger *slog.Logger ready atomic.Bool } func (is *IdentityService) GetPluginInfo(_ context.Context, _ *csiproto.GetPluginInfoRequest) (*csiproto.GetPluginInfoResponse, error) { return \u0026amp;csiproto.GetPluginInfoResponse{ Name: \u0026#34;csi.example.com\u0026#34;, VendorVersion: \u0026#34;1.0.0\u0026#34;, }, nil } func (is *IdentityService) GetPluginCapabilities(_ context.Context, _ *csiproto.GetPluginCapabilitiesRequest) (*csiproto.GetPluginCapabilitiesResponse, error) { return \u0026amp;csiproto.GetPluginCapabilitiesResponse{ Capabilities: []*csiproto.PluginCapability{ {Type: \u0026amp;csiproto.PluginCapability_Service_{ Service: \u0026amp;csiproto.PluginCapability_Service{ Type: csiproto.PluginCapability_Service_CONTROLLER_SERVICE, }, }}, }, }, nil } func (is *IdentityService) Probe(_ context.Context, _ *csiproto.ProbeRequest) (*csiproto.ProbeResponse, error) { return \u0026amp;csiproto.ProbeResponse{ Ready: \u0026amp;wrapperspb.BoolValue{Value: is.ready.Load()}, }, nil } Call is.SetReady(true) after everything is wired up, just before grpcServer.Serve.\nControllerService The controller manages your storage backend. The struct holds clients to your backend APIs:\ntype ControllerService struct { csiproto.UnimplementedControllerServer logger *slog.Logger volClient VolumeAPIClient vmClient VMAPIClient region string } CreateVolume must be idempotent — csi-provisioner will retry. Check by name first:\nfunc (cs *ControllerService) CreateVolume(ctx context.Context, req *csiproto.CreateVolumeRequest) (*csiproto.CreateVolumeResponse, error) { if req.GetName() == \u0026#34;\u0026#34; { return nil, status.Error(codes.InvalidArgument, \u0026#34;missing name\u0026#34;) } existing, err := cs.volClient.GetByName(ctx, req.GetName()) if err == nil { return buildCreateResponse(existing), nil } sizeGB := bytesToGB(req.GetCapacityRange().GetRequiredBytes()) vol, err := cs.volClient.Create(ctx, CreateVolumeParams{ Name: req.GetName(), SizeGB: sizeGB, AvailabilityZone: azFromTopology(req.GetAccessibilityRequirements()), }) if err != nil { return nil, status.Errorf(codes.Internal, \u0026#34;create volume failed: %v\u0026#34;, err) } err = cs.waitForStatus(ctx, vol.ID, \u0026#34;available\u0026#34;) if err != nil { return nil, status.Errorf(codes.Internal, \u0026#34;volume not ready: %v\u0026#34;, err) } return buildCreateResponse(vol), nil } The exponential backoff wait is critical — most storage APIs are asynchronous:\nfunc (cs *ControllerService) waitForStatus(ctx context.Context, volumeID, target string) error { backoff := wait.Backoff{ Duration: time.Second, Factor: 1.1, Steps: 15, } return wait.ExponentialBackoffWithContext(ctx, backoff, func(ctx context.Context) (bool, error) { vol, err := cs.volClient.Get(ctx, volumeID) if err != nil { return false, err } return vol.Status == target, nil }) } ControllerPublishVolume attaches the block device to the VM. The key output is PublishContext — this is how the device path travels from the controller to the node:\nfunc (cs *ControllerService) ControllerPublishVolume(ctx context.Context, req *csiproto.ControllerPublishVolumeRequest) (*csiproto.ControllerPublishVolumeResponse, error) { volumeID := req.GetVolumeId() nodeID := req.GetNodeId() err = cs.vmClient.AttachVolume(ctx, nodeID, volumeID) if err != nil { return nil, status.Errorf(codes.Internal, \u0026#34;attach failed: %v\u0026#34;, err) } err = cs.waitUntilAttached(ctx, nodeID, volumeID) if err != nil { return nil, status.Errorf(codes.Internal, \u0026#34;attach timeout: %v\u0026#34;, err) } devicePath, err := cs.volClient.GetDevicePath(ctx, volumeID, nodeID) if err != nil { return nil, status.Errorf(codes.Internal, \u0026#34;device path: %v\u0026#34;, err) } return \u0026amp;csiproto.ControllerPublishVolumeResponse{ PublishContext: map[string]string{ \u0026#34;DevicePath\u0026#34;: devicePath, }, }, nil } NodeService The node runs with elevated privileges on every machine. It receives the PublishContext from the controller and turns it into real mounts.\nNodeStageVolume formats (if needed) and mounts to a shared staging path:\nfunc (ns *NodeService) NodeStageVolume(ctx context.Context, req *csiproto.NodeStageVolumeRequest) (*csiproto.NodeStageVolumeResponse, error) { stagingTarget := req.GetStagingTargetPath() devicePath := req.GetPublishContext()[\u0026#34;DevicePath\u0026#34;] notMnt, err := ns.mounter.IsLikelyNotMountPointAttach(stagingTarget) if err != nil { return nil, status.Error(codes.Internal, err.Error()) } if notMnt { fsType := \u0026#34;ext4\u0026#34; if mnt := req.GetVolumeCapability().GetMount(); mnt != nil \u0026amp;\u0026amp; mnt.GetFsType() != \u0026#34;\u0026#34; { fsType = mnt.GetFsType() } err = ns.mounter.FormatAndMount(devicePath, stagingTarget, fsType, nil) if err != nil { return nil, status.Error(codes.Internal, err.Error()) } } return \u0026amp;csiproto.NodeStageVolumeResponse{}, nil } NodePublishVolume bind-mounts the staging path into the pod\u0026rsquo;s specific directory:\nfunc (ns *NodeService) NodePublishVolume(ctx context.Context, req *csiproto.NodePublishVolumeRequest) (*csiproto.NodePublishVolumeResponse, error) { targetPath := req.GetTargetPath() stagingPath := req.GetStagingTargetPath() mountOptions := []string{\u0026#34;bind\u0026#34;} if req.GetReadonly() { mountOptions = append(mountOptions, \u0026#34;ro\u0026#34;) } else { mountOptions = append(mountOptions, \u0026#34;rw\u0026#34;) } err := ns.mounter.Mount(stagingPath, targetPath, \u0026#34;ext4\u0026#34;, mountOptions) if err != nil { return nil, status.Errorf(codes.Internal, \u0026#34;bind mount failed: %v\u0026#34;, err) } return \u0026amp;csiproto.NodePublishVolumeResponse{}, nil } NodeGetInfo is called by kubelet at startup to learn the node\u0026rsquo;s identity and topology. This is how topology.kubernetes.io/zone gets set:\nfunc (ns *NodeService) NodeGetInfo(ctx context.Context, _ *csiproto.NodeGetInfoRequest) (*csiproto.NodeGetInfoResponse, error) { instanceID, err := ns.metadata.GetInstanceID() if err != nil { return nil, status.Errorf(codes.Internal, \u0026#34;instance ID: %v\u0026#34;, err) } zone, err := ns.metadata.GetAvailabilityZone() if err != nil { return nil, status.Errorf(codes.Internal, \u0026#34;availability zone: %v\u0026#34;, err) } return \u0026amp;csiproto.NodeGetInfoResponse{ NodeId: instanceID, AccessibleTopology: \u0026amp;csiproto.Topology{ Segments: map[string]string{ \u0026#34;topology.kubernetes.io/zone\u0026#34;: zone, }, }, }, nil } The Mounter Interface Rather than calling mount directly, wrap k8s.io/mount-utils. This makes the node service testable with a fake mounter:\ntype Mounter interface { FormatAndMount(source, target, fsType string, options []string) error Mount(source, target, fsType string, options []string) error UnmountPath(mountPath string) error IsLikelyNotMountPointAttach(path string) (bool, error) GetDevicePath(ctx context.Context, volumeID string) (string, error) Execer() exec.Interface } func NewMounter() Mounter { return \u0026amp;mounter{ BaseMounter: mountutil.New(\u0026#34;\u0026#34;), exec: exec.New(), } } Deployment Controller Deployment apiVersion: apps/v1 kind: Deployment metadata: name: csi-controller spec: replicas: 1 template: spec: containers: - name: csi-provisioner image: registry.k8s.io/sig-storage/csi-provisioner:v5.2.0 args: [\u0026#34;--csi-address=/csi/csi.sock\u0026#34;, \u0026#34;--leader-election\u0026#34;] volumeMounts: - name: socket-dir mountPath: /csi - name: csi-attacher image: registry.k8s.io/sig-storage/csi-attacher:v4.8.0 args: [\u0026#34;--csi-address=/csi/csi.sock\u0026#34;, \u0026#34;--leader-election\u0026#34;] volumeMounts: - name: socket-dir mountPath: /csi - name: csi-driver image: your-registry/csi-controller:latest env: - name: CSI_ENDPOINT value: unix:///csi/csi.sock volumeMounts: - name: socket-dir mountPath: /csi volumes: - name: socket-dir emptyDir: {} # shared between all sidecars in this pod Node DaemonSet apiVersion: apps/v1 kind: DaemonSet metadata: name: csi-node spec: template: spec: containers: - name: node-driver-registrar image: registry.k8s.io/sig-storage/csi-node-driver-registrar:v2.13.0 args: - --csi-address=/csi/csi.sock - --kubelet-registration-path=/var/lib/kubelet/plugins/csi.example.com/csi.sock volumeMounts: - name: socket-dir mountPath: /csi - name: registration-dir mountPath: /registration - name: csi-driver image: your-registry/csi-node:latest securityContext: privileged: true capabilities: add: [\u0026#34;SYS_ADMIN\u0026#34;] env: - name: CSI_ENDPOINT value: unix:///csi/csi.sock volumeMounts: - name: socket-dir mountPath: /csi - name: kubelet-dir mountPath: /var/lib/kubelet mountPropagation: Bidirectional # essential: propagates mounts to host - name: dev-dir mountPath: /dev volumes: - name: socket-dir hostPath: path: /var/lib/kubelet/plugins/csi.example.com/ type: DirectoryOrCreate - name: registration-dir hostPath: path: /var/lib/kubelet/plugins_registry/ - name: kubelet-dir hostPath: path: /var/lib/kubelet - name: dev-dir hostPath: path: /dev Two things that catch people out:\nmountPropagation: Bidirectional on the kubelet dir — without this, mounts made inside the node container aren\u0026rsquo;t visible to the host, so the pod never sees the volume. privileged: true — the node driver calls mount(2), which requires root. There\u0026rsquo;s no way around it. Error Handling and Idempotency The CSI spec requires every RPC to be idempotent. Kubernetes retries. You will receive CreateVolume twice with the same name. You will receive ControllerPublishVolume for an already-attached volume. Handle this by checking state before acting:\n// CreateVolume — return existing volume if found by name. existing, err := cs.volClient.GetByName(ctx, name) if err == nil { return buildCreateResponse(existing), nil } // ControllerPublishVolume — skip attach if already attached. if attached, _ := cs.isAttached(ctx, nodeID, volumeID); attached { devicePath, _ := cs.getDevicePath(ctx, volumeID, nodeID) return \u0026amp;csiproto.ControllerPublishVolumeResponse{ PublishContext: map[string]string{\u0026#34;DevicePath\u0026#34;: devicePath}, }, nil } Use gRPC status codes correctly:\ncodes.InvalidArgument // caller\u0026#39;s fault, missing/bad input codes.NotFound // resource doesn\u0026#39;t exist codes.AlreadyExists // for CreateVolume when name conflicts with different params codes.Internal // backend/driver error codes.Unavailable // transient, safe to retry Node Registration Sequence sequenceDiagram participant REG as node-driver-registrar participant DRV as NodeService participant KBL as kubelet REG-\u003e\u003eDRV: GetPluginInfo() → {name, version} REG-\u003e\u003eDRV: GetPluginCapabilities() REG-\u003e\u003eKBL: register at /var/lib/kubelet/plugins_registry/csi.example.com-reg.sock KBL-\u003e\u003eREG: NodeGetInfo request REG-\u003e\u003eDRV: NodeGetInfo() → {nodeID, topology} REG--\u003e\u003eKBL: {nodeID, driverSocket=/var/lib/kubelet/plugins/csi.example.com/csi.sock} Note over KBL: kubelet routes volume opsfor this driver to the node socket KBL-\u003e\u003eDRV: NodeStageVolume / NodePublishVolume / ... sequenceDiagram participant REG as node-driver-registrar participant DRV as NodeService participant KBL as kubelet REG-\u003e\u003eDRV: GetPluginInfo() → {name, version} REG-\u003e\u003eDRV: GetPluginCapabilities() REG-\u003e\u003eKBL: register at /var/lib/kubelet/plugins_registry/csi.example.com-reg.sock KBL-\u003e\u003eREG: NodeGetInfo request REG-\u003e\u003eDRV: NodeGetInfo() → {nodeID, topology} REG--\u003e\u003eKBL: {nodeID, driverSocket=/var/lib/kubelet/plugins/csi.example.com/csi.sock} Note over KBL: kubelet routes volume opsfor this driver to the node socket KBL-\u003e\u003eDRV: NodeStageVolume / NodePublishVolume / ... ⤢ Common Pitfalls Volume size rounding. Many backends work in GiB, not bytes. req.GetCapacityRange().GetRequiredBytes() is in bytes. Convert carefully and round up, not down — returning a smaller volume than requested violates the spec.\nNodeStageVolume vs NodePublishVolume. Stage runs once per volume per node. Publish runs once per pod. If you skip staging and try to format in publish, two pods requesting the same volume will race to run mkfs on the same device.\nBlocking RPCs. All CSI RPCs are synchronous from the caller\u0026rsquo;s perspective, but backends are async. CreateVolume must not return until the volume is usable. Poll with backoff; don\u0026rsquo;t sleep for a fixed duration.\nxfs and duplicate UUIDs. If you restore a snapshot to create a new volume, the new volume has the same filesystem UUID as the original. xfs refuses to mount two volumes with the same UUID on the same node. Pass nouuid as a mount option for xfs volumes:\nif fsType == \u0026#34;xfs\u0026#34; { options = append(options, \u0026#34;nouuid\u0026#34;) } Block volumes. If you support VolumeCapability_AccessMode_SINGLE_NODE_WRITER for raw block, NodePublishVolume needs a different code path — create a file at targetPath and bind-mount the device file, not a directory.\nTesting Unit test the controller and node services with mock backends and a fake mounter:\nfunc TestCreateVolume_Idempotent(t *testing.T) { ctrl := NewControllerService(fakeVolClient, fakeVMClient, ...) req := \u0026amp;csiproto.CreateVolumeRequest{ Name: \u0026#34;test-vol\u0026#34;, CapacityRange: \u0026amp;csiproto.CapacityRange{RequiredBytes: 16 * 1024 * 1024 * 1024}, VolumeCapabilities: defaultCaps(), } resp1, err := ctrl.CreateVolume(ctx, req) require.NoError(t, err) resp2, err := ctrl.CreateVolume(ctx, req) require.NoError(t, err) assert.Equal(t, resp1.GetVolume().GetVolumeId(), resp2.GetVolume().GetVolumeId()) } For the node service, k8s.io/mount-utils ships a FakeMounter that records mount calls without touching the filesystem — use it.\nIntegration testing the full flow requires a real Kubernetes cluster. The Kubernetes CSI sanity test suite provides a standard set of conformance tests you can run against a live driver socket.\nWhat You End Up With After all of this, what you have is:\nA controller binary that speaks to your storage API and manages volume/snapshot lifecycle A node binary that runs on every node, formats disks, and manages mounts Two Kubernetes workloads (Deployment + DaemonSet) with the right sidecar containers StorageClass, RBAC, and a driver registration object pointing at your plugin name The CSI spec is verbose but consistent. Once you\u0026rsquo;ve implemented CreateVolume and NodeStageVolume, the pattern repeats. The hard parts are operational: idempotency under retries, async backend polling, and getting mount propagation right in the DaemonSet.\nThe official CSI spec and the kubernetes-csi examples repo are the two references worth keeping open while building. Note, even though the examples repo says this is not a good example, it\u0026rsquo;s actually a decent start to get the boilerplate and setup right.\n","permalink":"https://syndbg.dev/posts/2026-04-24-writing-a-kubernetes-csi-driver/","summary":"\u003cp\u003eKubernetes storage is one of those areas that looks simple from the outside — you create a \u003ccode\u003ePersistentVolumeClaim\u003c/code\u003e, a pod mounts it, done. But the moment you need to integrate your own storage backend, you\u0026rsquo;re staring at the \u003ca href=\"https://github.com/container-storage-interface/spec\"\u003eContainer Storage Interface spec\u003c/a\u003e, sidecar containers you\u0026rsquo;ve never heard of, and gRPC services that have to be wired together just right.\u003c/p\u003e\n\u003cp\u003eI had the pleasure of writing and contributing to a production-grade CSI driver end-to-end. This post covers everything I wish I had in one place: what CSI actually is, how Kubernetes orchestrates it, and how to implement all three services in Go.\u003c/p\u003e","title":"Writing a Kubernetes CSI Driver: Controller and Node from Scratch"},{"content":"Go 1.25 just dropped with expected changes to GOMAXPROCS, which significantly change how Go applications behave in containerized environments. The runtime now automatically detects and respects container CPU limits when setting GOMAXPROCS. This isn\u0026rsquo;t just a minor improvement—it\u0026rsquo;s a shift that may dramatically improve performance for millions of containerized Go applications.\nAnd this isn\u0026rsquo;t the only amazing change, but this is the one I\u0026rsquo;ll focus on in this post.\nThe Problem That Plagued Go for Years Before Go 1.25, there was a fundamental mismatch between Go\u0026rsquo;s runtime and containerized environments:\n# Your Kubernetes pod with 2 CPU cores resources: limits: cpu: \u0026#34;2\u0026#34; # But Go saw the entire host machine runtime.NumCPU() // Returns 8 (host CPUs) runtime.GOMAXPROCS(0) // Also 8 - ignoring your 2-core limit! This led to:\nOver-scheduling: Go created 8 goroutines for CPU-bound work on a 2-core container Context switching overhead: Excessive goroutine switching degraded performance Resource contention: Multiple containers fighting for the same CPU cores Manual workarounds: Developers had to manually set GOMAXPROCS everywhere Popular libraries like uber-go/automaxprocs existed solely to fix this fundamental issue, showing how widespread the problem was.\nWhat Changed in Go 1.25 Go 1.25 introduces two features:\n1. Container-Aware GOMAXPROCS (Linux) The runtime now reads cgroup CPU bandwidth limits and automatically sets GOMAXPROCS accordingly:\n// On a container with 2 CPU cores limit: runtime.NumCPU() // Still returns 8 (host CPUs) runtime.GOMAXPROCS(0) // Now returns 2 (respects container limit!) 2. Dynamic GOMAXPROCS Updates The runtime periodically updates GOMAXPROCS if CPU availability changes:\n// Your application automatically adapts to: // - Kubernetes horizontal pod autoscaling // - Container CPU limit changes // - Node CPU availability changes Seeing It in Action Let me show you the difference with a practical example. Here\u0026rsquo;s a test program that reveals the new behavior:\npackage main import ( \u0026#34;fmt\u0026#34; \u0026#34;runtime\u0026#34; \u0026#34;time\u0026#34; ) func main() { fmt.Printf(\u0026#34;Go version: %s\\n\u0026#34;, runtime.Version()) fmt.Printf(\u0026#34;Host CPUs: %d\\n\u0026#34;, runtime.NumCPU()) fmt.Printf(\u0026#34;GOMAXPROCS: %d\\n\u0026#34;, runtime.GOMAXPROCS(0)) // Monitor for changes (Go 1.25 feature) for i := 0; i \u0026lt; 10; i++ { time.Sleep(5 * time.Second) fmt.Printf(\u0026#34;Time %ds: GOMAXPROCS = %d\\n\u0026#34;, (i+1)*5, runtime.GOMAXPROCS(0)) } } Kubernetes Results Comparison Go 1.24 and earlier:\n# Pod with 2 CPU cores limit resources: limits: cpu: \u0026#34;2\u0026#34; Go version: go1.24.0 Host CPUs: 8 GOMAXPROCS: 8 # ❌ Ignores container limit Go 1.25:\nGo version: go1.25.0 Host CPUs: 8 GOMAXPROCS: 2 # ✅ Respects container limit! Docker Results Running the same program in Docker with CPU limits:\n# Docker with 1.5 CPU cores docker run --cpus=\u0026#34;1.5\u0026#34; golang:1.25 go run main.go Go version: go1.25.0 Host CPUs: 8 GOMAXPROCS: 2 # ✅ Rounds up fractional CPU limits Complete Examples and Testing I\u0026rsquo;ve created comprehensive examples to test all scenarios. Here are the key ones:\nKubernetes Manifest Examples # Example 1: CPU-limited container apiVersion: apps/v1 kind: Deployment metadata: name: go-app-cpu-limited spec: template: spec: containers: - name: go-app image: golang:1.25 resources: limits: cpu: \u0026#34;2\u0026#34; # GOMAXPROCS will be 2 memory: \u0026#34;256Mi\u0026#34; command: [\u0026#34;go\u0026#34;, \u0026#34;run\u0026#34;, \u0026#34;/app/main.go\u0026#34;] # Example 2: Manual override (disables auto-detection) apiVersion: apps/v1 kind: Deployment metadata: name: go-app-manual spec: template: spec: containers: - name: go-app image: golang:1.25 resources: limits: cpu: \u0026#34;2\u0026#34; # CPU limit is 2 env: - name: GOMAXPROCS value: \u0026#34;8\u0026#34; # Manual override - will use 8 command: [\u0026#34;go\u0026#34;, \u0026#34;run\u0026#34;, \u0026#34;/app/main.go\u0026#34;] When the Magic Doesn\u0026rsquo;t Happen The automatic behavior is disabled in these cases:\n1. Manual GOMAXPROCS Setting // Environment variable GOMAXPROCS=8 go run main.go // Or in code runtime.GOMAXPROCS(8) 2. GODEBUG Overrides # Disable container awareness GODEBUG=containermaxprocs=0 go run main.go # Disable dynamic updates GODEBUG=updatemaxprocs=0 go run main.go # Disable both GODEBUG=containermaxprocs=0,updatemaxprocs=0 go run main.go 3. Non-Linux Platforms Currently, container-aware GOMAXPROCS only works on Linux (where cgroups exist).\nMigration Guide for Existing Applications 1. Remove Manual GOMAXPROCS Workarounds Before (Go \u0026lt; 1.25):\nimport ( _ \u0026#34;go.uber.org/automaxprocs\u0026#34; // Popular workaround library ) func init() { // Manual calculation based on container limits if cpuLimit := os.Getenv(\u0026#34;CPU_LIMIT\u0026#34;); cpuLimit != \u0026#34;\u0026#34; { if limit, err := strconv.Atoi(cpuLimit); err == nil { runtime.GOMAXPROCS(limit) } } } After (Go 1.25):\n// Just delete all the manual workarounds! // Go handles it automatically now 2. Test Thoroughly Since this changes fundamental runtime behavior, comprehensive testing is critical:\n# Test your application with various CPU limits docker run --cpus=\u0026#34;1\u0026#34; myapp:go1.25 docker run --cpus=\u0026#34;2\u0026#34; myapp:go1.25 docker run --cpus=\u0026#34;4\u0026#34; myapp:go1.25 # Monitor performance metrics kubectl top pods 3. Monitor Key Metrics Watch these metrics after upgrading:\nCPU utilization efficiency Response times Context switch rates Memory usage patterns Goroutine counts Cloud Platform Compatibility This change affects major cloud platforms:\nAWS ECS Fargate: ✅ Respects CPU limits EKS: ✅ Works with pod CPU limits Lambda: ✅ May improve performance in custom runtimes Google Cloud Cloud Run: ✅ Automatically optimizes for allocated CPU GKE: ✅ Full Kubernetes support Cloud Functions: ✅ Benefits custom Go runtimes Azure Container Instances: ✅ Respects CPU limits AKS: ✅ Full Kubernetes support Functions: ✅ Helps custom Go runtimes Performance Tuning Considerations When Automatic is Perfect Standard web applications: REST APIs, microservices Data processing pipelines: ETL jobs, stream processing Background workers: Queue processors, schedulers When Manual Tuning May Be Better CPU-intensive algorithms: May need fine-tuning for specific workloads Memory-bound applications: GOMAXPROCS might not be the bottleneck Legacy applications: With complex custom scheduling logic Testing Your Specific Workload func benchmarkWithDifferentGOMAXPROCS() { for _, procs := range []int{1, 2, 4, 8} { runtime.GOMAXPROCS(procs) start := time.Now() // Your workload here runYourWorkload() fmt.Printf(\u0026#34;GOMAXPROCS=%d: %v\\n\u0026#34;, procs, time.Since(start)) } } Troubleshooting Common Issues Issue 1: Performance Regression Symptoms: Application slower after upgrading to Go 1.25 Solution:\n# Temporarily disable to confirm GODEBUG=containermaxprocs=0 go run main.go # Or use manual setting GOMAXPROCS=8 go run main.go Issue 2: GOMAXPROCS Not Changing Check: Container actually has CPU limits set\n# In container, check cgroup limits cat /sys/fs/cgroup/cpu.max cat /sys/fs/cgroup/cpu/cpu.cfs_quota_us Issue 3: Fractional CPU Confusion Remember: Go rounds UP fractional CPU limits\nCPU Limit: 1.5 → GOMAXPROCS: 2 CPU Limit: 2.7 → GOMAXPROCS: 3 CPU Limit: 3.9 → GOMAXPROCS: 4 The Bigger Picture: Why This Matters This change represents a fundamental shift in how Go applications integrate with modern infrastructure:\n1. Cloud-Native by Default Go applications now understand their containerized environment without additional configuration.\n2. Better Resource Efficiency Automatic optimization means better cost efficiency in cloud deployments.\n3. Simplified Operations DevOps teams no longer need to manually tune GOMAXPROCS for every deployment.\n4. Platform Consistency The same application binary adapts to different deployment environments automatically.\nFuture Implications This opens doors for even smarter runtime optimizations:\nMemory-aware garbage collection tuning Network buffer sizing based on container limits Dynamic goroutine pool sizing Automatic backpressure mechanisms Conclusion Go 1.25\u0026rsquo;s container-aware GOMAXPROCS is more than just a feature—it\u0026rsquo;s a fundamental improvement that makes Go applications truly cloud-native by default. The days of manual GOMAXPROCS tuning and third-party workaround libraries are finally over.\nKey takeaways:\n✅ Automatic container CPU limit detection on Linux ✅ Dynamic updates when limits change ✅ Zero configuration required ✅ Significant performance improvements in many cases ✅ Backward compatible with manual overrides If you\u0026rsquo;re running Go applications in containers (and who isn\u0026rsquo;t these days?), Go 1.25 should be at the top of your upgrade priority list. Test it thoroughly, but expect to see measurable performance improvements across your containerized Go workloads.\nWant to test this yourself? All the examples and code from this post are included below. Try them out in your own Kubernetes cluster or Docker environment to see the magic in action!\nComplete Test Code Examples Kubernetes Manifests Here are the complete Kubernetes manifests for testing different scenarios:\n# kubernetes-manifests.yaml --- # Example 1: CPU-limited container that will respect limits apiVersion: apps/v1 kind: Deployment metadata: name: go-app-cpu-limited-1core labels: app: go-gomaxprocs-test scenario: cpu-limited-1core spec: replicas: 1 selector: matchLabels: app: go-gomaxprocs-test scenario: cpu-limited-1core template: metadata: labels: app: go-gomaxprocs-test scenario: cpu-limited-1core spec: containers: - name: go-app image: golang:1.25 resources: limits: cpu: \u0026#34;1\u0026#34; # GOMAXPROCS should be 1 memory: \u0026#34;256Mi\u0026#34; requests: cpu: \u0026#34;500m\u0026#34; memory: \u0026#34;128Mi\u0026#34; command: [\u0026#34;sh\u0026#34;, \u0026#34;-c\u0026#34;] args: - | cat \u0026gt; /tmp/main.go \u0026lt;\u0026lt; \u0026#39;EOF\u0026#39; package main import ( \u0026#34;fmt\u0026#34; \u0026#34;runtime\u0026#34; \u0026#34;time\u0026#34; ) func main() { fmt.Printf(\u0026#34;=== Go 1.25 GOMAXPROCS Test (1 CPU limit) ===\\n\u0026#34;) fmt.Printf(\u0026#34;Go version: %s\\n\u0026#34;, runtime.Version()) fmt.Printf(\u0026#34;Host CPUs: %d\\n\u0026#34;, runtime.NumCPU()) fmt.Printf(\u0026#34;GOMAXPROCS: %d\\n\u0026#34;, runtime.GOMAXPROCS(0)) fmt.Printf(\u0026#34;Expected GOMAXPROCS: 1 (CPU limit)\\n\\n\u0026#34;) for i := 0; i \u0026lt; 6; i++ { time.Sleep(10 * time.Second) fmt.Printf(\u0026#34;Time %ds: GOMAXPROCS = %d\\n\u0026#34;, (i+1)*10, runtime.GOMAXPROCS(0)) } } EOF cd /tmp \u0026amp;\u0026amp; go run main.go --- # Example 2: CPU-limited container with 2 cores apiVersion: apps/v1 kind: Deployment metadata: name: go-app-cpu-limited-2core labels: app: go-gomaxprocs-test scenario: cpu-limited-2core spec: replicas: 1 selector: matchLabels: app: go-gomaxprocs-test scenario: cpu-limited-2core template: metadata: labels: app: go-gomaxprocs-test scenario: cpu-limited-2core spec: containers: - name: go-app image: golang:1.25 resources: limits: cpu: \u0026#34;2\u0026#34; # GOMAXPROCS should be 2 memory: \u0026#34;512Mi\u0026#34; requests: cpu: \u0026#34;1\u0026#34; memory: \u0026#34;256Mi\u0026#34; command: [\u0026#34;sh\u0026#34;, \u0026#34;-c\u0026#34;] args: - | cat \u0026gt; /tmp/main.go \u0026lt;\u0026lt; \u0026#39;EOF\u0026#39; package main import ( \u0026#34;fmt\u0026#34; \u0026#34;runtime\u0026#34; \u0026#34;time\u0026#34; ) func main() { fmt.Printf(\u0026#34;=== Go 1.25 GOMAXPROCS Test (2 CPU limit) ===\\n\u0026#34;) fmt.Printf(\u0026#34;Go version: %s\\n\u0026#34;, runtime.Version()) fmt.Printf(\u0026#34;Host CPUs: %d\\n\u0026#34;, runtime.NumCPU()) fmt.Printf(\u0026#34;GOMAXPROCS: %d\\n\u0026#34;, runtime.GOMAXPROCS(0)) fmt.Printf(\u0026#34;Expected GOMAXPROCS: 2 (CPU limit)\\n\\n\u0026#34;) for i := 0; i \u0026lt; 6; i++ { time.Sleep(10 * time.Second) fmt.Printf(\u0026#34;Time %ds: GOMAXPROCS = %d\\n\u0026#34;, (i+1)*10, runtime.GOMAXPROCS(0)) } } EOF cd /tmp \u0026amp;\u0026amp; go run main.go --- # Example 3: Manual GOMAXPROCS override (should ignore container limits) apiVersion: apps/v1 kind: Deployment metadata: name: go-app-manual-override labels: app: go-gomaxprocs-test scenario: manual-override spec: replicas: 1 selector: matchLabels: app: go-gomaxprocs-test scenario: manual-override template: metadata: labels: app: go-gomaxprocs-test scenario: manual-override spec: containers: - name: go-app image: golang:1.25 resources: limits: cpu: \u0026#34;1\u0026#34; # CPU limit is 1 memory: \u0026#34;256Mi\u0026#34; env: - name: GOMAXPROCS value: \u0026#34;4\u0026#34; # Manual override - should use 4, not 1 command: [\u0026#34;sh\u0026#34;, \u0026#34;-c\u0026#34;] args: - | cat \u0026gt; /tmp/main.go \u0026lt;\u0026lt; \u0026#39;EOF\u0026#39; package main import ( \u0026#34;fmt\u0026#34; \u0026#34;runtime\u0026#34; \u0026#34;time\u0026#34; ) func main() { fmt.Printf(\u0026#34;=== Go 1.25 GOMAXPROCS Test (Manual Override) ===\\n\u0026#34;) fmt.Printf(\u0026#34;Go version: %s\\n\u0026#34;, runtime.Version()) fmt.Printf(\u0026#34;Host CPUs: %d\\n\u0026#34;, runtime.NumCPU()) fmt.Printf(\u0026#34;GOMAXPROCS: %d\\n\u0026#34;, runtime.GOMAXPROCS(0)) fmt.Printf(\u0026#34;CPU Limit: 1, Manual GOMAXPROCS: 4\\n\u0026#34;) fmt.Printf(\u0026#34;Expected GOMAXPROCS: 4 (manual override)\\n\\n\u0026#34;) for i := 0; i \u0026lt; 6; i++ { time.Sleep(10 * time.Second) fmt.Printf(\u0026#34;Time %ds: GOMAXPROCS = %d\\n\u0026#34;, (i+1)*10, runtime.GOMAXPROCS(0)) } } EOF cd /tmp \u0026amp;\u0026amp; go run main.go --- # Example 4: Fractional CPU limits (should round up) apiVersion: apps/v1 kind: Deployment metadata: name: go-app-fractional-cpu labels: app: go-gomaxprocs-test scenario: fractional-cpu spec: replicas: 1 selector: matchLabels: app: go-gomaxprocs-test scenario: fractional-cpu template: metadata: labels: app: go-gomaxprocs-test scenario: fractional-cpu spec: containers: - name: go-app image: golang:1.25 resources: limits: cpu: \u0026#34;1.5\u0026#34; # GOMAXPROCS should be 2 (rounds up) memory: \u0026#34;256Mi\u0026#34; requests: cpu: \u0026#34;750m\u0026#34; memory: \u0026#34;128Mi\u0026#34; command: [\u0026#34;sh\u0026#34;, \u0026#34;-c\u0026#34;] args: - | cat \u0026gt; /tmp/main.go \u0026lt;\u0026lt; \u0026#39;EOF\u0026#39; package main import ( \u0026#34;fmt\u0026#34; \u0026#34;runtime\u0026#34; \u0026#34;time\u0026#34; ) func main() { fmt.Printf(\u0026#34;=== Go 1.25 GOMAXPROCS Test (1.5 CPU limit) ===\\n\u0026#34;) fmt.Printf(\u0026#34;Go version: %s\\n\u0026#34;, runtime.Version()) fmt.Printf(\u0026#34;Host CPUs: %d\\n\u0026#34;, runtime.NumCPU()) fmt.Printf(\u0026#34;GOMAXPROCS: %d\\n\u0026#34;, runtime.GOMAXPROCS(0)) fmt.Printf(\u0026#34;Expected GOMAXPROCS: 2 (1.5 rounds up to 2)\\n\\n\u0026#34;) for i := 0; i \u0026lt; 6; i++ { time.Sleep(10 * time.Second) fmt.Printf(\u0026#34;Time %ds: GOMAXPROCS = %d\\n\u0026#34;, (i+1)*10, runtime.GOMAXPROCS(0)) } } EOF cd /tmp \u0026amp;\u0026amp; go run main.go Docker Compose Test Scenarios Complete Docker Compose setup for testing different CPU limits:\n# docker-compose.yaml version: \u0026#39;3.8\u0026#39; services: # Test 1: 1 CPU limit go-app-1cpu: image: golang:1.25 container_name: go-1.25-test-1cpu deploy: resources: limits: cpus: \u0026#39;1.0\u0026#39; memory: 256M reservations: cpus: \u0026#39;0.5\u0026#39; memory: 128M command: \u0026gt; sh -c \u0026#34; cat \u0026gt; /tmp/test.go \u0026lt;\u0026lt; \u0026#39;EOF\u0026#39; package main import ( \\\u0026#34;fmt\\\u0026#34; \\\u0026#34;runtime\\\u0026#34; \\\u0026#34;time\\\u0026#34; ) func main() { fmt.Printf(\\\u0026#34;=== Docker Test: 1.0 CPU Limit ===\\n\\\u0026#34;) fmt.Printf(\\\u0026#34;Go version: %s\\n\\\u0026#34;, runtime.Version()) fmt.Printf(\\\u0026#34;Host CPUs: %d\\n\\\u0026#34;, runtime.NumCPU()) fmt.Printf(\\\u0026#34;GOMAXPROCS: %d\\n\\\u0026#34;, runtime.GOMAXPROCS(0)) fmt.Printf(\\\u0026#34;Expected: 1\\n\\n\\\u0026#34;) for i := 0; i \u0026lt; 12; i++ { time.Sleep(5 * time.Second) fmt.Printf(\\\u0026#34;[%02d] GOMAXPROCS = %d\\n\\\u0026#34;, i+1, runtime.GOMAXPROCS(0)) } } EOF cd /tmp \u0026amp;\u0026amp; go run test.go \u0026#34; # Test 2: 2.5 CPU limit (should round up to 3) go-app-2-5cpu: image: golang:1.25 container_name: go-1.25-test-2-5cpu deploy: resources: limits: cpus: \u0026#39;2.5\u0026#39; memory: 512M reservations: cpus: \u0026#39;1.0\u0026#39; memory: 256M command: \u0026gt; sh -c \u0026#34; cat \u0026gt; /tmp/test.go \u0026lt;\u0026lt; \u0026#39;EOF\u0026#39; package main import ( \\\u0026#34;fmt\\\u0026#34; \\\u0026#34;runtime\\\u0026#34; \\\u0026#34;time\\\u0026#34; ) func main() { fmt.Printf(\\\u0026#34;=== Docker Test: 2.5 CPU Limit ===\\n\\\u0026#34;) fmt.Printf(\\\u0026#34;Go version: %s\\n\\\u0026#34;, runtime.Version()) fmt.Printf(\\\u0026#34;Host CPUs: %d\\n\\\u0026#34;, runtime.NumCPU()) fmt.Printf(\\\u0026#34;GOMAXPROCS: %d\\n\\\u0026#34;, runtime.GOMAXPROCS(0)) fmt.Printf(\\\u0026#34;Expected: 3 (2.5 rounds up)\\n\\n\\\u0026#34;) for i := 0; i \u0026lt; 12; i++ { time.Sleep(5 * time.Second) fmt.Printf(\\\u0026#34;[%02d] GOMAXPROCS = %d\\n\\\u0026#34;, i+1, runtime.GOMAXPROCS(0)) } } EOF cd /tmp \u0026amp;\u0026amp; go run test.go \u0026#34; # Test 3: Manual override go-app-manual: image: golang:1.25 container_name: go-1.25-test-manual environment: - GOMAXPROCS=8 deploy: resources: limits: cpus: \u0026#39;1.0\u0026#39; memory: 256M command: \u0026gt; sh -c \u0026#34; cat \u0026gt; /tmp/test.go \u0026lt;\u0026lt; \u0026#39;EOF\u0026#39; package main import ( \\\u0026#34;fmt\\\u0026#34; \\\u0026#34;runtime\\\u0026#34; \\\u0026#34;time\\\u0026#34; ) func main() { fmt.Printf(\\\u0026#34;=== Docker Test: Manual Override ===\\n\\\u0026#34;) fmt.Printf(\\\u0026#34;Go version: %s\\n\\\u0026#34;, runtime.Version()) fmt.Printf(\\\u0026#34;Host CPUs: %d\\n\\\u0026#34;, runtime.NumCPU()) fmt.Printf(\\\u0026#34;GOMAXPROCS: %d\\n\\\u0026#34;, runtime.GOMAXPROCS(0)) fmt.Printf(\\\u0026#34;CPU Limit: 1.0, GOMAXPROCS env: 8\\n\\\u0026#34;) fmt.Printf(\\\u0026#34;Expected: 8 (manual override)\\n\\n\\\u0026#34;) for i := 0; i \u0026lt; 12; i++ { time.Sleep(5 * time.Second) fmt.Printf(\\\u0026#34;[%02d] GOMAXPROCS = %d\\n\\\u0026#34;, i+1, runtime.GOMAXPROCS(0)) } } EOF cd /tmp \u0026amp;\u0026amp; go run test.go \u0026#34; # Test 4: GODEBUG disable go-app-godebug: image: golang:1.25 container_name: go-1.25-test-godebug environment: - GODEBUG=containermaxprocs=0 deploy: resources: limits: cpus: \u0026#39;1.0\u0026#39; memory: 256M command: \u0026gt; sh -c \u0026#34; cat \u0026gt; /tmp/test.go \u0026lt;\u0026lt; \u0026#39;EOF\u0026#39; package main import ( \\\u0026#34;fmt\\\u0026#34; \\\u0026#34;runtime\\\u0026#34; \\\u0026#34;time\\\u0026#34; ) func main() { fmt.Printf(\\\u0026#34;=== Docker Test: GODEBUG Disabled ===\\n\\\u0026#34;) fmt.Printf(\\\u0026#34;Go version: %s\\n\\\u0026#34;, runtime.Version()) fmt.Printf(\\\u0026#34;Host CPUs: %d\\n\\\u0026#34;, runtime.NumCPU()) fmt.Printf(\\\u0026#34;GOMAXPROCS: %d\\n\\\u0026#34;, runtime.GOMAXPROCS(0)) fmt.Printf(\\\u0026#34;CPU Limit: 1.0, GODEBUG=containermaxprocs=0\\n\\\u0026#34;) fmt.Printf(\\\u0026#34;Expected: %d (same as host CPUs)\\n\\n\\\u0026#34;, runtime.NumCPU()) for i := 0; i \u0026lt; 12; i++ { time.Sleep(5 * time.Second) fmt.Printf(\\\u0026#34;[%02d] GOMAXPROCS = %d\\n\\\u0026#34;, i+1, runtime.GOMAXPROCS(0)) } } EOF cd /tmp \u0026amp;\u0026amp; go run test.go \u0026#34; Interactive Test Program An interactive program to test GOMAXPROCS behavior in different scenarios:\n// gomaxprocs-test.go package main import ( \u0026#34;fmt\u0026#34; \u0026#34;os\u0026#34; \u0026#34;runtime\u0026#34; \u0026#34;strconv\u0026#34; \u0026#34;time\u0026#34; ) func main() { fmt.Println(\u0026#34;🚀 Go 1.25 Container-Aware GOMAXPROCS Interactive Test\u0026#34;) fmt.Println(\u0026#34;=====================================================\u0026#34;) // Display current environment showEnvironmentInfo() // Test scenarios fmt.Println(\u0026#34;\\n🧪 Running test scenarios...\u0026#34;) testScenario1_DefaultBehavior() testScenario2_ManualOverride() testScenario3_MonitorChanges() } func showEnvironmentInfo() { fmt.Printf(\u0026#34;\\n📊 Environment Information:\\n\u0026#34;) fmt.Printf(\u0026#34; Go Version: %s\\n\u0026#34;, runtime.Version()) fmt.Printf(\u0026#34; GOOS: %s\\n\u0026#34;, runtime.GOOS) fmt.Printf(\u0026#34; GOARCH: %s\\n\u0026#34;, runtime.GOARCH) fmt.Printf(\u0026#34; Host CPUs: %d\\n\u0026#34;, runtime.NumCPU()) fmt.Printf(\u0026#34; Current GOMAXPROCS: %d\\n\u0026#34;, runtime.GOMAXPROCS(0)) // Check environment variables if gomaxprocs := os.Getenv(\u0026#34;GOMAXPROCS\u0026#34;); gomaxprocs != \u0026#34;\u0026#34; { fmt.Printf(\u0026#34; GOMAXPROCS env: %s\\n\u0026#34;, gomaxprocs) } if godebug := os.Getenv(\u0026#34;GODEBUG\u0026#34;); godebug != \u0026#34;\u0026#34; { fmt.Printf(\u0026#34; GODEBUG: %s\\n\u0026#34;, godebug) } // Check cgroup information (Linux only) if runtime.GOOS == \u0026#34;linux\u0026#34; { checkCgroupInfo() } } func checkCgroupInfo() { fmt.Printf(\u0026#34;\\n🐧 Linux cgroup Information:\\n\u0026#34;) // Check cgroup v2 CPU limits if data, err := os.ReadFile(\u0026#34;/sys/fs/cgroup/cpu.max\u0026#34;); err == nil { fmt.Printf(\u0026#34; cpu.max: %s\u0026#34;, string(data)) } // Check cgroup v1 CPU limits if data, err := os.ReadFile(\u0026#34;/sys/fs/cgroup/cpu/cpu.cfs_quota_us\u0026#34;); err == nil { fmt.Printf(\u0026#34; cfs_quota_us: %s\u0026#34;, string(data)) } if data, err := os.ReadFile(\u0026#34;/sys/fs/cgroup/cpu/cpu.cfs_period_us\u0026#34;); err == nil { fmt.Printf(\u0026#34; cfs_period_us: %s\u0026#34;, string(data)) } } func testScenario1_DefaultBehavior() { fmt.Printf(\u0026#34;\\n🔍 Test 1: Default Behavior\\n\u0026#34;) fmt.Printf(\u0026#34; Expected: GOMAXPROCS should respect container CPU limits\\n\u0026#34;) original := runtime.GOMAXPROCS(0) fmt.Printf(\u0026#34; Result: GOMAXPROCS = %d\\n\u0026#34;, original) if runtime.GOOS == \u0026#34;linux\u0026#34; { if original \u0026lt;= runtime.NumCPU() { fmt.Printf(\u0026#34; ✅ Looks good! GOMAXPROCS ≤ Host CPUs\\n\u0026#34;) } else { fmt.Printf(\u0026#34; ❓ Unexpected: GOMAXPROCS \u0026gt; Host CPUs\\n\u0026#34;) } } else { fmt.Printf(\u0026#34; ℹ️ Container awareness only works on Linux\\n\u0026#34;) } } func testScenario2_ManualOverride() { fmt.Printf(\u0026#34;\\n🔧 Test 2: Manual Override\\n\u0026#34;) fmt.Printf(\u0026#34; Testing manual GOMAXPROCS setting...\\n\u0026#34;) original := runtime.GOMAXPROCS(0) testValue := original + 2 fmt.Printf(\u0026#34; Setting GOMAXPROCS to %d\\n\u0026#34;, testValue) runtime.GOMAXPROCS(testValue) actual := runtime.GOMAXPROCS(0) fmt.Printf(\u0026#34; Result: GOMAXPROCS = %d\\n\u0026#34;, actual) if actual == testValue { fmt.Printf(\u0026#34; ✅ Manual override works correctly\\n\u0026#34;) } else { fmt.Printf(\u0026#34; ❌ Manual override failed\\n\u0026#34;) } // Restore original runtime.GOMAXPROCS(original) fmt.Printf(\u0026#34; Restored to original value: %d\\n\u0026#34;, original) } func testScenario3_MonitorChanges() { fmt.Printf(\u0026#34;\\n⏰ Test 3: Monitor for Dynamic Changes\\n\u0026#34;) fmt.Printf(\u0026#34; Monitoring GOMAXPROCS for 30 seconds...\\n\u0026#34;) fmt.Printf(\u0026#34; (In Go 1.25, this should update if container limits change)\\n\\n\u0026#34;) for i := 0; i \u0026lt; 6; i++ { time.Sleep(5 * time.Second) gomaxprocs := runtime.GOMAXPROCS(0) fmt.Printf(\u0026#34; [%02ds] GOMAXPROCS = %d\\n\u0026#34;, (i+1)*5, gomaxprocs) } fmt.Printf(\u0026#34;\\n ℹ️ In a real container environment with changing limits,\\n\u0026#34;) fmt.Printf(\u0026#34; you might see GOMAXPROCS change dynamically.\\n\u0026#34;) } Usage Instructions Running the Kubernetes Tests # Apply all manifests kubectl apply -f kubernetes-manifests.yaml # Check the results kubectl logs deployment/go-app-cpu-limited-1core kubectl logs deployment/go-app-cpu-limited-2core kubectl logs deployment/go-app-manual-override kubectl logs deployment/go-app-fractional-cpu # Clean up kubectl delete -f kubernetes-manifests.yaml Running the Docker Compose Tests # Run all test scenarios docker-compose up # Run individual tests docker-compose up go-app-1cpu docker-compose up go-app-2-5cpu docker-compose up go-app-manual docker-compose up go-app-godebug # Clean up docker-compose down Running the Interactive Tests # Save as gomaxprocs-test.go and run directly go run gomaxprocs-test.go # Or in a container with CPU limits docker run --cpus=\u0026#34;2\u0026#34; golang:1.25 sh -c \u0026#34; cat \u0026gt; test.go \u0026lt;\u0026lt; \u0026#39;EOF\u0026#39; [paste the gomaxprocs-test.go code here] EOF go run test.go \u0026#34; Related Resources Go 1.25 Release Notes Container-aware GOMAXPROCS Documentation cgroups Documentation ","permalink":"https://syndbg.dev/posts/2025-01-13-go-1-25/","summary":"\u003cp\u003eGo 1.25 just dropped with expected changes to GOMAXPROCS, which significantly change how Go applications behave in containerized environments. The runtime now \u003cstrong\u003eautomatically detects and respects container CPU limits\u003c/strong\u003e when setting \u003ccode\u003eGOMAXPROCS\u003c/code\u003e. This isn\u0026rsquo;t just a minor improvement—it\u0026rsquo;s a shift that \u003cstrong\u003emay\u003c/strong\u003e dramatically improve performance for millions of containerized Go applications.\u003c/p\u003e\n\u003cp\u003eAnd this isn\u0026rsquo;t the only amazing change, but this is the one I\u0026rsquo;ll focus on in this post.\u003c/p\u003e\n\u003ch2 id=\"the-problem-that-plagued-go-for-years\"\u003eThe Problem That Plagued Go for Years\u003c/h2\u003e\n\u003cp\u003eBefore Go 1.25, there was a fundamental mismatch between Go\u0026rsquo;s runtime and containerized environments:\u003c/p\u003e","title":"Practically, Go 1.25's Container-Aware GOMAXPROCS: What You Need to Know"},{"content":"Anton Antonov Engineering Manager and Principal Software Engineer, focused on developer experience, platform engineering, and distributed systems.\nGo practitioner for many years, building distributed systems, sometimes with Raft consensus. Kubernetes practitioner focused on cluster ops, multi-cloud, and security (cert-manager, PKI automation). Nowadays also writing production Rust. I\u0026rsquo;ve built closed-source platform systems and open-sourced pieces of them along the way — see projects.\nBy preference: Go, Rust, CockroachDB, Redis, Terraform, and whatever cloud environment the problem actually needs.\nBased in Sofia, Bulgaria.\nContact GitHub: github.com/syndbg LinkedIn: linkedin.com/in/syndbg ","permalink":"https://syndbg.dev/about/","summary":"\u003ch1 id=\"anton-antonov\"\u003eAnton Antonov\u003c/h1\u003e\n\u003cp\u003eEngineering Manager and Principal Software Engineer, focused on developer experience, platform engineering, and distributed systems.\u003c/p\u003e\n\u003cp\u003eGo practitioner for many years, building distributed systems, sometimes with Raft consensus. Kubernetes practitioner focused on cluster ops, multi-cloud, and security (cert-manager, PKI automation). Nowadays also writing production Rust. I\u0026rsquo;ve built closed-source platform systems and open-sourced pieces of them along the way — see \u003ca href=\"/projects/\"\u003eprojects\u003c/a\u003e.\u003c/p\u003e\n\u003cp\u003eBy preference: Go, Rust, CockroachDB, Redis, Terraform, and whatever cloud environment the problem actually needs.\u003c/p\u003e","title":"About"},{"content":"","permalink":"https://syndbg.dev/llms.txt","summary":"","title":"llms.txt"}]