Skip to content

Routing series: Post 1 taxonomy (publish) + Post 2 cluster aliasing (draft) - #289

Open
SamBarker wants to merge 22 commits into
kroxylicious:mainfrom
SamBarker:routing-series/post2-cluster-aliasing
Open

SamBarker wants to merge 22 commits into
kroxylicious:mainfrom
SamBarker:routing-series/post2-cluster-aliasing

Conversation

@SamBarker

Copy link
Copy Markdown
Member

Post 1 — "These are not the brokers you are connecting to" — introduces the Routing API, establishes the series vocabulary, and names the four routing patterns. Scheduled for publication Friday 10 October.

Post 2 — "Cluster Aliasing: Identity without commitment" — provides the technical depth for the Current SF session: what aliasing is, how the proxy makes it work (node IDs, topology discovery, PID management), and where aliasing ends. Left as a draft pending a publication date.

Replaces #287 (closed).

@SamBarker
SamBarker requested a review from a team as a code owner October 6, 2026 01:26
Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
- Clarify MetadataResponse structure (broker list + partition→nodeId mapping)
- Tighten node ID collision explanation: proxy cannot have two entries for node ID 2
- Reframe bijection as per-DAG-node with downstream/upstream ID spaces
- Add N upstream ID spaces explanation — the reason a bijection is necessary
- Replace LaTeX formula (doesn't render) with plain English arithmetic
- Make diagram bidirectional and show fan-out to multiple clusters
- Topology discovery: add 'no independent view' framing and per-router
  shared cache explanation with additive semantics safety argument
- PID section: lead with the multi-cluster collision problem, add honest
  note that translation table is in-flight (topic partition routing work)

Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
- Ground readers with an explanation of idempotent delivery and PIDs
  before introducing the multi-cluster problem
- Note idempotent delivery is default (and non-optional in Kafka 4.0)
- Clarify the collision is between brokers on different clusters
- Explain why a translation table rather than a bijection: PIDs are
  opaque broker-assigned integers with no exploitable arithmetic
  structure, unlike node IDs
- Per-router scoping: mirrors the node ID bijection structure with
  downstream virtual PIDs mapping to upstream broker-allocated PIDs
- Replace 'physical PID' with 'broker-allocated' throughout
- Explain why session-scoped is sufficient: brokers own transactional
  PID continuity, proxy rebuilds downstream-to-upstream mapping on
  each new session

Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
@SamBarker
SamBarker force-pushed the routing-series/post2-cluster-aliasing branch from f2b3cbc to 54b9589 Compare October 6, 2026 01:28
Comment thread _docs Outdated
Comment thread results.tar.gz Outdated
Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
@SamBarker
SamBarker force-pushed the routing-series/post2-cluster-aliasing branch from 54b9589 to 8ccc4f9 Compare October 6, 2026 01:33

@robobario robobario left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM thanks @SamBarker

Comment thread _posts/2026-10-10-these-are-not-the-brokers-you-are-connecting-to.md Outdated
Comment thread _posts/2026-10-10-these-are-not-the-brokers-you-are-connecting-to.md Outdated
SamBarker and others added 2 commits October 6, 2026 16:42
Co-authored-by: Robert Young <robertyoungnz@gmail.com>
Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
Co-authored-by: Robert Young <robertyoungnz@gmail.com>
Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
Comment thread _posts/2026-10-10-these-are-not-the-brokers-you-are-connecting-to.md Outdated
Comment thread _posts/2026-10-10-these-are-not-the-brokers-you-are-connecting-to.md Outdated
Comment thread _posts/2026-10-10-these-are-not-the-brokers-you-are-connecting-to.md Outdated
Comment thread _posts/2026-10-10-these-are-not-the-brokers-you-are-connecting-to.md Outdated
…n, restructure topic weaving

- Rename 'cluster aliasing' -> 'connection switching' in post 1 and post 2
- Rename 'union clusters' -> 'topic weaving' in post 1
- Rename 'virtual topics' -> 'partition weaving' in post 1
- Add stream branching as a new pattern section in post 1
- Fix connection switching closing paragraph (T3 resolution)
- Restructure topic weaving: Keith's NY/london/tokyo example leads, George Street becomes the collision case (T4 resolution)
- Remove forward post number commitments; use pattern names instead
- Update post 2 title and all body references to connection switching

Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
…tructure

- Introduce all four pattern names upfront in the intro paragraph
- Replace DAG composability paragraph with single-axis protocol depth framing
- Fix stale 'virtual topic' and 'alias' references in DAG paragraph
- Restructure topic weaving: Keith's NY/london/tokyo example leads, George
  Street becomes the collision case
- Tighten transaction boundary warning: 'no way to make it otherwise'
- Replace exhaustive RPC list with 'they get everywhere, as you might expect'
- Fix 'believes' typo; remove redundant broker-control sentence

Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
…en copy

- Fix broken post_url tag in intro
- Replace '2am nightmare' post 1 callback with self-contained opener
- Remove production shadowing section (stream branching pattern, not connection switching)
- Replace 'no stitching' with 'no weaving' to use established taxonomy
- Stash PID section to stash-replica-routing.md; add concise PID notes under 'Proxies aren't magic'
- New 'When is 2 != 2?' heading for node ID collision section
- Various copy fixes: spinning rust, vendor-neutral replication tool list, Bugatti callback, overclaiming on cutover promise

Signed-off-by: Sam Barker <sam@quadrocket.co.uk>

Jumping up to Layer 7 changes the game. Rather than blindly punting TCP bytes around, the proxy speaks Kafka. It knows a `Metadata` response from a `Produce` request, and it can answer a client's question about where the brokers live with whatever answer you need.

From day one, Kroxylicious gave you a Virtual Kafka Cluster (VKC) — a stable endpoint for your clients. But until recently, it was an opinionated pipe. That was a compliment: it did real work. It could encrypt records, enforce auth, rewrite headers — all without the client noticing. But it was still a fixed pipe: one client socket in, one backend socket out, wired to whatever cluster you pointed it at on boot. The filters were stateless with respect to the broker topology: an encryption filter might talk out-of-band to a KMS for keys, but it didn't care which physical cluster lived behind the socket.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: client socket goes to a broker, not a cluster.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in df2f68b — updated to broker rather than cluster.


## Name them we must

After reviewing the [routing design proposal](https://github.com/kroxylicious/design/pull/70) and thinking through the problem space, I think there are four distinct deployment patterns worth naming: **connection switching**, **stream branching**, **topic weaving**, and **partition weaving**. As with all patterns the boundaries are fuzzy and often more than one applies at once. The thing that convinces me they're real is that they have descriptive power — each one has its own tradeoffs, failure modes, and a distinct place on the spectrum from the clear(ish) waters of connection switching to the mangrove swamp of partition weaving. Like all good abstractions, they're useful.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like connection switching. I understood topic weaving and partiton weaving right away. stream branching had me wondering what was coming.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Glad the others landed. Stream branching was the best I could come up with — open to better suggestions. Branching streams is another candidate; the difference is whether you're naming the act (branching a stream) or the result (streams that branch). I went with the former but I'm not attached to it.


If you're reaching for the [Enterprise Integration Patterns](https://www.enterpriseintegrationpatterns.com/patterns/messaging/MessageRoutingIntro.html) book right now — well done for being as old as I am, the rest of you just Googled it. It describes what happens to individual messages, which is interesting, but I'm talking about whole streams. Those patterns are what makes all this possible under the hood; what I'm describing sits a layer or two higher.

The patterns do build on each other along a single axis: how deep into the Kafka protocol the router has to reach. Connection switching operates at the connection level — which broker, which forms part of which cluster, is this connection actually being sent to? Stream branching works at the message level — where does a copy of this record go? Topic weaving reaches into the topic catalog — which topics are visible and where do they live? Partition weaving goes deepest — how do physical partition spaces map onto a logical one? Each step requires the router(s) to understand more of the protocol than the last. But the complexity doesn't compound as routing is managed through a DAG which allows them to be composed rather than monolithic. A router handling topic visibility doesn't care whether the partitions behind it are woven from multiple clusters. A router rewriting partition numbers doesn't care whether the topic catalog above it spans two physical clusters or twenty.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"Stream branching works at the message level — where does a copy of this record go" - remember that Kafka's unit of currency is a batch rather than an individual record. I think we've shied away from routing records - at least for now, although it remains an worth aspiration.,

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point — you're right that we don't want to focus on record-level routing here. The design space does cover it (a router reconstituting a branch request can go as fine-grained as it likes), but that's not where current implementations sit. I've tweaked the language in the axis paragraph to avoid the ambiguity, and the stream branching section now covers the routing granularity point in more detail.


You're chasing an issue in the reporting pipeline that only shows up in the sixth hour of the run and nobody can figure out which record trips it up. You've been there, right? What you really want is access to the live data with a debugger. Alas, this job is stuffed full of Personally Identifiable Information — so you're out of luck. You've been there and got *that* t-shirt. What now? You build a router that shadows production traffic to a staging cluster, piping it through a redaction filter on the way so the PII becomes gibberish before it touches staging brokers. The client is still producing to one VKC. But the proxy is now opening *two* backend connections — one to the primary cluster, one to staging — and writing every record to both. The acknowledgement the client gets back is from the primary; the shadow write happens on the side. One stream in, two streams out. If you've got your EIP book off the shelf already, yes — it's a Wire Tap with a Message Translator on the diverted stream.

That's one flavour of stream branching — the secondary stream is a copy of the primary, diverted somewhere and the client never realises. But the diverted branch doesn't have to be a copy. Every record passes through the routing DAG before it hits the broker. Something in that DAG can inspect the payload, compute something from it — consumer lag, record counts by key, an audit trail your compliance team needs — and emit that derived data to a topic anywhere the routing config points — a different topic on the same cluster, or a cluster the client has never heard of. The primary stream is untouched. The additional streams just appear, derived from the original client's activity. The clients don't change. Which means anything that previously required every application to cooperate — the audit trail, the telemetry, the thing all but one of your clients never supported — you can now just invent that stream yourself.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this part less convincing. I keep thinking - why's wouldn't I just use Kafka Streams?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Horses for courses — Kafka Streams is absolutely the right answer for some of this. Where the proxy approach has the edge: no duplicate storage cost (you're deriving data in flight rather than materialising a new topic), and no framework overhead for what are often stateless transforms. Kafka Streams and Connect are powerful but they're not lightweight — if all you need is "emit a record count to a metrics topic as this batch passes through", standing up a Streams application is a lot of machinery for a stateless operation. I've added a sentence to the section to make this clearer.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thinking about the awkward audience member with an awkward question is a good strategy :)


That's the common case: different topics, different clusters, one coherent view for the client. Your router is what holds the map — which topic lives where, which broker list to return, what gets dispatched where.

There's a subtler case. What if two of your clusters both host a topic called `orders`? Think of it like a sorting office: you write *George Street* on the parcel and drop it at the counter. The sorting office decides whether you mean George Street, Edinburgh or George Street, Dunedin — a city in New Zealand that Scottish settlers named after Edinburgh, gave the same street names, and promptly scrambled the layout. (Guess why I know.) The sender doesn't need to know both exist. Neither does your client. Because the proxy controls what gets returned in a metadata response, your router decides which topics are visible and under what names — which means in principle it can arbitrate which names reach clients, or whether the duplicate surfaces at all.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ahhhh! I get you now!


There's a subtler case. What if two of your clusters both host a topic called `orders`? Think of it like a sorting office: you write *George Street* on the parcel and drop it at the counter. The sorting office decides whether you mean George Street, Edinburgh or George Street, Dunedin — a city in New Zealand that Scottish settlers named after Edinburgh, gave the same street names, and promptly scrambled the layout. (Guess why I know.) The sender doesn't need to know both exist. Neither does your client. Because the proxy controls what gets returned in a metadata response, your router decides which topics are visible and under what names — which means in principle it can arbitrate which names reach clients, or whether the duplicate surfaces at all.

You look after a gaggle of Kafka clusters. You know how they got there. You're not proud of all of them. You can't herd them — geese are worse than cats, they honk back — but you don't have to admit to anyone else how many geese there actually are, or what they're called.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is a good point. System often grow organically. The proxy lets you hide this complexity from the clients - and gives you the cluster admin freedom, to incrementally deal with the mess.


A single logical topic whose partitions are composed from physical topics on multiple clusters — which don't even need to share a name. A client asks for metadata for `topic-x` and gets back 32 partitions; partitions 0–15 are sourced from `topic-x` on `us-east`, 16–31 from `topic-eu` on `eu-west`. A `ProduceRequest` for partition 20 gets routed to `eu-west` transparently, with partition numbers rewritten to match the physical layout. The client sees one topic. It has no idea.

The motivating example here is one topic weaving can't solve. Your reporting pipeline in `us-east` is hardcoded to read from `topic-x`. It has always read from `topic-x`. It will continue to read from `topic-x`. The problem is on the producer side: your business is growing in Europe, and sending every EU event across the Atlantic to land on the `us-east` cluster is expensive, slow, and fragile. The EU team provisions `topic-eu` on a cluster local to them, sized for their producer load. The router presents both physical topics as one logical `topic-x` to every client. EU producers write locally. The reporting pipeline reads the full partition space without a config change. Nobody crosses the Atlantic unnecessarily.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's making the producers in the EU choose to write to partitions in the EU. Kafka clients choose which partition to send data, so what will stop it sending my batch to partition 0 (in the US).

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oops that wasn't quite the example I was trying to make. I've tweaked the text to make it clearer the woven topic is only on the consumer side.

@k-wall k-wall left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi @SamBarker, I've only made it part way through again. Sorry today got rather swallowed up by other things.

Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
You're chasing an issue in the reporting pipeline that only shows up in the sixth hour of the run and nobody can figure out which record trips it up. You've been there, right? What you really want is access to the live data with a debugger. Alas, this job is stuffed full of Personally Identifiable Information — so you're out of luck. You've been there and got *that* t-shirt. What now? You build a router that shadows production traffic to a staging cluster, piping it through a redaction filter on the way so the PII becomes gibberish before it touches staging brokers. The client is still producing to one VKC. But the proxy is now opening *two* backend connections — one to the primary cluster, one to staging — and writing every record to both. The acknowledgement the client gets back is from the primary; the shadow write happens on the side. One stream in, two streams out. If you've got your EIP book off the shelf already, yes — it's a Wire Tap with a Message Translator on the diverted stream.

That's one flavour of stream branching — the secondary stream is a copy of the primary, diverted somewhere and the client never realises. But the diverted branch doesn't have to be a copy. Every record passes through the routing DAG before it hits the broker. Something in that DAG can inspect the payload, compute something from it — consumer lag, record counts by key, an audit trail your compliance team needs — and emit that derived data to a topic anywhere the routing config points — a different topic on the same cluster, or a cluster the client has never heard of. The primary stream is untouched. The additional streams just appear, derived from the original client's activity. The clients don't change. Which means anything that previously required every application to cooperate — the audit trail, the telemetry, the thing all but one of your clients never supported — you can now just invent that stream yourself.
That's one flavour of stream branching — the secondary stream is a copy of the primary, diverted somewhere and the client never realises. But the diverted branch doesn't have to be a copy. Every record passes through the routing DAG before it hits the broker. Something in that DAG can inspect the payload, compute something from it — consumer lag, record counts by key, an audit trail your compliance team needs — and emit that derived data to a topic anywhere the routing config points — a different topic on the same cluster, or a cluster the client has never heard of. The primary stream is untouched. The additional streams just appear, derived from the original client's activity. The clients don't change. Which means anything that previously required every application to cooperate — the audit trail, the telemetry, the thing all but one of your clients never supported — you can now just invent that stream yourself. This all works for stateless transformations performed inline. If you need state, joins, or windowing, Kafka Streams or Kafka Connect have your back.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's better, but I've still got questions:

I'm not sure about the consumer lag idea. Anything observing the fetches to compute lag would rely on the fact that a consumer was actually alive in some form. A dead consumer would generate no fetches, so I would have no 'lag signal'. Observing consumer lag on the cluster is the only way to get a complete picture, isn't it?

Record accounts by key - yeah - that works to some extent. But if you've scaled your proxy replicas, each replica sees only part of the picture. What aggregates the aggregations?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair point on consumer lag — you're right that cluster-side observation is the proper place for complete lag metrics. On counts/telemetry, the proxy isn't doing the final aggregation — it's just emitting inline record telemetry for downstream systems to aggregate. I've updated the examples in this section to focus on clean, stateless inline operations (a redaction filter for PII, a sampling filter to thin traffic for smaller staging environments, and derivation filters for audit logs/telemetry).

A single logical topic whose partitions are composed from physical topics on multiple clusters — which don't even need to share a name. A client asks for metadata for `topic-x` and gets back 32 partitions; partitions 0–15 are sourced from `topic-x` on `us-east`, 16–31 from `topic-eu` on `eu-west`. A `ProduceRequest` for partition 20 gets routed to `eu-west` transparently, with partition numbers rewritten to match the physical layout. The client sees one topic. It has no idea.

The motivating example here is one topic weaving can't solve. Your reporting pipeline in `us-east` is hardcoded to read from `topic-x`. It has always read from `topic-x`. It will continue to read from `topic-x`. The problem is on the producer side: your business is growing in Europe, and sending every EU event across the Atlantic to land on the `us-east` cluster is expensive, slow, and fragile. The EU team provisions `topic-eu` on a cluster local to them, sized for their producer load. The router presents both physical topics as one logical `topic-x` to every client. EU producers write locally. The reporting pipeline reads the full partition space without a config change. Nobody crosses the Atlantic unnecessarily.
The motivating example here is one topic weaving can't solve. Your reporting pipeline in `us-east` is hardcoded to read from `topic-x`. It has always read from `topic-x`. It will continue to read from `topic-x`. The problem is on the producer side: your business is growing in Europe, and sending every EU event across the Atlantic to land on the `us-east` cluster is expensive, slow, and fragile. The EU team provisions `topic-eu` on a cluster local to them, sized for their producer load. EU producers write to `topic-eu` on that cluster, proxies optional. The router presents both physical topics as one logical `topic-x` to the reporting pipeline. It reads the full partition space without a config change. Nobody crosses the Atlantic unnecessarily.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok I see the point now. You use the proxy to hide some complexity from reporting application. Rather than it having a connection to two clusters and reading from both topics, it now reads from topic-x which is partition woven into one. It is a nice simply facade. I get that.

If you didn't have partition weaving but only topic weaving, you could get the same result by reporting app consuming both topic-x and topic-eu (I mean Consumer#assign(Collection var1)). The reporting app gets a single stream of records which in reality are coming from two separate clusters.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, exactly — if you control the consumer and don't mind it managing multiple topic subscriptions/assignments, topic weaving gets you part of the way there. Partition weaving is what lets you keep that complexity entirely in the infrastructure so downstream consumers don't even need to know the second topic exists.

I've reworked this section to frame it around extending weaving down to Kafka's fundamental unit of parallelism (partitions), swapped the example to a third-party risk engine consuming an orders topic, and explained how consumer group coordination remains pinned to a single cluster without the coordinator needing to know about the seam.

- Introduce partition weaving as extending weaving to Kafka's unit of parallelism
- Switch illustrative example from topic-x to orders with third-party risk engine
- Clarify consumer group coordinator mechanics and offset tracking across clusters

Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
… railway metaphor

Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
…and consumer group RPC demuxing

Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
Signed-off-by: Sam Barker <sam@quadrocket.co.uk>
Comment thread _drafts/routing-post2.md

And this isn't just about inspecting frames at Layer 7. If all you have is a filter pipeline on a static 1:1 pipe, understanding that a `Fetch` belongs in `eu-west-1a` doesn't help you—the pipe is already nailed to broker 1 in `eu-west-1c`. The breakthrough comes from the Routing API making the backend connection dynamic.

When a `Fetch` request arrives at a Kroxylicious instance in `eu-west-1a` — on the pipe for the leader over in `eu-west-1c` — the router inspects the partition's in-sync replica (ISR) set. If broker 3 is sitting right next door in `eu-west-1a` with a caught-up replica, the router doesn't rewrite the frame—it dynamically dispatches that individual request down a separate backend pipe to broker 3:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This got me thinking about OffsetCommit and where that needs to go. I think the answer is it always goes to the leader, even if a follower-fetch scenario. The OffsetCommit carries the leaderEpoch which will have come from the follower fetched response. I think this idea works.

Comment thread _drafts/routing-post2.md

And this isn't just about inspecting frames at Layer 7. If all you have is a filter pipeline on a static 1:1 pipe, understanding that a `Fetch` belongs in `eu-west-1a` doesn't help you—the pipe is already nailed to broker 1 in `eu-west-1c`. The breakthrough comes from the Routing API making the backend connection dynamic.

When a `Fetch` request arrives at a Kroxylicious instance in `eu-west-1a` — on the pipe for the leader over in `eu-west-1c` — the router inspects the partition's in-sync replica (ISR) set. If broker 3 is sitting right next door in `eu-west-1a` with a caught-up replica, the router doesn't rewrite the frame—it dynamically dispatches that individual request down a separate backend pipe to broker 3:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's a wrinkle - you probably don't need to consider it because the idea is sound. Fetches can address more than one topic, so it is possible that follower isn't the appropiate destination for all the topic-partition targets. The fetch might need to be decomposed.

Comment thread _drafts/routing-post2.md

## How the proxy makes it work

Making one physical cluster look like a different one at Layer 7 requires more than TCP forwarding. There are three places where the proxy has to do actual work.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This section gets quite dense quite quickly. I think you are do with a paragraph the introduces the problems that needs to be solved first at a high level, then let the sections get into the nitty gritty.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants