Skip to main content
Experimental feature. Response caching is in alpha and is subject to change, and should not be used in production environments.
The router caches two kinds of subgraph response. Entity fetches. An entity fetch is the request the router sends to a subgraph through the _entities root field to resolve fields of an entity that another subgraph owns. Entries are keyed per entity and per selection set, so asking for different fields of the same entity does not share an entry. Root query fetches. A root query fetch is the request the router sends a subgraph for the root fields the client selected. The whole answer is stored as one entry. The key covers the operation text and every variable value. Any difference in either is a miss, and the request goes to the subgraph as it would have. Only queries are cached. Mutations and subscriptions are never cached. Response caching is disabled by default.

What gets cached

A subgraph response is cached only when its Cache-Control header asks to be. The rules apply in this order:
  1. no-store refuses caching.
  2. no-cache refuses caching, in any form. This includes the qualified no-cache="Set-Cookie" form.
  3. private marks the response as belonging to a private context. It is cached only under a key scoped to that context, and only when private_id is configured and resolves for the request. See Private responses. Without one, private refuses caching.
  4. s-maxage sets the lifetime and takes precedence over max-age.
  5. max-age sets the lifetime when s-maxage is absent.
  6. When neither lifetime is present, fallback_ttl applies if the header contains public, private, must-revalidate, proxy-revalidate, stale-if-error, or stale-while-revalidate. The stale directives require a valid duration, such as stale-if-error=60.
public is optional. For example, Cache-Control: max-age=60 and Cache-Control: s-maxage=60 both allow caching for 60 seconds. A value of 0 for the selected lifetime refuses caching. It does not fall through to the next rule. For example, max-age=60, s-maxage=0 is not cached, while max-age=0, s-maxage=60 is cached for 60 seconds. An invalid duration in a recognized directive also prevents caching. A missing or empty Cache-Control header is never cached. Headers containing only unrecognized directives, such as cdn-cache-control=60 or immutable, are also never cached. fallback_ttl does not apply to these responses. Revalidation and stale-response directives opt into the fallback lifetime. The response cache does not revalidate entries or serve expired entries. A response carrying GraphQL errors is never cached, whatever its Cache-Control header says.
Configuring the response cache does not make anything cacheable on its own. Subgraphs opt in with a recognized caching directive, such as Cache-Control: max-age=60 or Cache-Control: public. If nothing appears to be cached, check the subgraph response headers first.

Request directives

A client controls the cache for its own request through the Cache-Control request header. Two directives are honored:
  • no-cache is never answered from the cache. Every fetch of the operation goes to its subgraph, and the responses are stored as they are for any other request, so a no-cache request refreshes the entries it bypassed.
  • no-store keeps the responses to this request out of the cache. Nothing is stored, but a warm entry still answers the request, as RFC 9111 allows.
Sent together, no-cache, no-store neither reads nor writes the cache. Every other request directive is ignored, including max-age, min-fresh, max-stale and only-if-cached, as is the Pragma header. A Cache-Control request header that does not parse is ignored as a whole.
Any client can send these directives, and they cannot be turned off. A client that sends no-cache on every request bypasses the cache on every request, and its subgraphs see all of its traffic. A no-cache request is not deduplicated with identical requests in flight either.

Enable it

Redis

Every router replica shares one cache and entries outlive the process. Define a storage provider, then reference it by provider_id. Redis 7.0 or newer is required to use this feature.
An unknown provider_id fails startup.

Router memory

Each replica holds its own cache, nothing survives a restart, and no Redis is needed.
provider defaults to redis, so caching in memory has to be asked for by name. This prevents a missing provider_id from quietly turning one shared cache into one cache per replica.

Configuration reference

config.yaml
storage is required when enabled is true. Environment variables bypass the config schema validation. In each enabled scope, a fallback_ttl of zero or less fails startup, as does a private_id that does not compile.

Per-subgraph configuration

all applies to every subgraph. An entry under subgraphs, keyed by subgraph name, replaces all for that subgraph:
config.yaml
Here inventory is never cached, products keeps responses without a lifetime for five minutes, and every other subgraph keeps them for thirty seconds. An entry is a full override and is explicit. Nothing under all reaches a subgraph that has an entry, and no key is defaulted:
  • enabled must be set. An entry without it does not cache that subgraph. A fallback_ttl on its own does not enable anything.
  • An enabled entry must set fallback_ttl, or startup fails naming the entry.
  • An entry without private_id has none. So products above does not cache its private responses, although all names an id. Repeat the private_id in the entry to keep it.
Setting all.enabled to false turns the cache off for every subgraph without an entry, so only the entries that set enabled: true are cached. With no such entry the router starts and logs a warning, and nothing is cached. An entry naming a subgraph the graph does not have is kept and logged at startup, since a feature flag graph may have it. The invalidation indexes, the invalidation endpoint, and the cache tag header are not per subgraph.

Private responses

A subgraph whose answer depends on who is asking marks it Cache-Control: private. A shared cache must not serve that answer to anyone else, but the router can still cache it for the context it was produced in, so the same context asking again is served without a subgraph request. To do that the router needs an id for that context. private_id is an expression evaluated once per request that yields it, such as a JWT claim or a request header. The id partitions the cache key: what it names is up to the expression, and it need not be a single user. A tenant, an organisation, or a role are all valid contexts, as long as every request that shares the id may share the response.
config.yaml
Under all it applies to every subgraph. Under a subgraphs entry it applies to that subgraph alone, and an entry without one has none. See Per-subgraph configuration. With a private_id in place:
  • A private response is stored under a key that includes the id, so each context has an entry of its own. A lookup with one id never finds an entry stored under another.
  • A request for which the expression yields nothing, such as an anonymous request whose claim is nil, is not cached when the subgraph answers private. The subgraph is asked every time. This is not an error.
  • public responses are unaffected. They are stored under the shared key and served to every request, whatever the id.
  • With cache_control_policy or a most_restrictive_cache_control response header rule in place, a response served from a private entry reaches the client with Cache-Control: private and the entry’s remaining lifetime, so a browser or CDN in front of the router does not share it either. Without one, the router sends no Cache-Control for the hit.
  • Invalidation by subgraph, type, or cache tag drops private entries like any other.
The id is digested with SHA-256 before it becomes part of the key, so raw ids never reach the store. The expression must evaluate to a string. A value of any other type is reported through the response cache warning described under Observability and treated as no id.
private_id partitions the cache. It does not by itself guarantee that a response is served only to those it was meant for: two requests with the same id share an entry, whatever else differs between them. Propagated headers are not part of the key unless the subgraph names them in Vary, so a private answer that also depends on a header the id does not account for is served to every request with that id. Name the header in Vary, fold it into the expression, or have the subgraph answer no-store.An expression that yields the same value for different contexts, or one that can be set by the client itself, lets one context be served another’s entry. Prefer a verified claim over a plain header.A set response header rule on Cache-Control replaces what the subgraph sent, private included. A hit served from a private entry is treated the same as the origin response, so such a rule can hand the client a public policy for it.

Varying responses

A subgraph whose answer depends on a request header, such as Accept-Language, names it in a Vary response header. The router then keeps one entry per set of values it sent, learned from the first response, the way an HTTP proxy cache does:
  • The first request for an operation misses. The names in the response’s Vary header are stored at the entry’s key as a record, and the body is stored under a variant key derived from the values the router sent.
  • A later lookup finds the record, digests the values this request sends the subgraph for those headers, and looks the variant up under the resulting key. A varying response costs two cache round trips per lookup. A response without Vary costs one, as before.
  • Values are taken from the request as the subgraph receives it, after header propagation. A header named in Vary that the router does not forward counts as empty, so it never splits the cache.
  • Vary: * is never cached, and neither is a response naming more than 100 headers.
  • A header whose value differs per caller, such as Cookie or Authorization, yields one variant per caller. Nobody is served another caller’s body, but the entries are rarely reused. A subgraph whose answer is per caller should say private instead.
  • A subgraph that changes what it varies on adds the new set of names to the record on its next miss, newest first, rather than replacing it. A lookup tries one variant per set in that order and serves the first it finds, so variants stored under the old names stay reachable until they expire. A record keeps at most 8 sets; past that the oldest are dropped.
  • With engine.enable_multi_fetch on, a merged entity fetch carries one Vary for the whole response, and the router applies it to every alias in it. An alias whose own answer would have varied on less is stored under the wider set, which splits its entries by headers they do not depend on. That costs extra misses, never a wrong body. A merged fetch follows records the same way a single fetch does, with one second round trip for all of its aliases together. The lever for the cost is on the subgraph side: emit the minimum Vary for what was actually resolved rather than a fixed union for the service.

Limitations

These apply to the current alpha. A batch is served from the cache only when every entity in it is present. One miss sends the whole batch to the subgraph, including the entities that were already cached. Entities are stored individually, so a later batch made up only of cached entities is a full hit. Forwarded request headers reach the cache key only through Vary. The key covers the request the router renders for the subgraph, which is the operation text and its variables. Headers added by header propagation are applied after that and never reach the key on their own. Two requests that differ only by a propagated header share one entry unless the subgraph names that header in Vary, as described under Varying responses. A subgraph that varies by request header without saying so must use no-store, or private together with a private_id whose value accounts for that header. Omitting public does not prevent caching. Cache hits produce no subgraph telemetry. A hit skips the subgraph fetch, so no subgraph span or metric is recorded for it. Subgraph request counts fall as the hit rate rises. Use them to confirm the cache is working, not to measure traffic. Merged entity fetches have their own rules. With engine.enable_multi_fetch on, entity fetches to the same subgraph within one parallel wave are merged into a single request with aliased _entities fields. The response is taken apart again on the way into the cache, so each alias stores one entry per entity, the same as the unmerged fetch it replaces would have. From there:
  • Each alias is served all-or-nothing, like a batch. A warm alias is left out of the next merged request. A cold one is fetched whole.
  • Only a request that finds every alias warm is a cache hit. When some aliases are warm and others are not, the router still sends the request for the cold ones.
  • The reported Cache-Control is capped at the warm aliases. The lifetime reported for a partly warm fetch is the life left on its warm aliases.
  • Declared cache tags are dropped. apolloEntityCacheTags is one flat list with nothing to say which alias each entry belongs to. The entries are still indexed by subgraph and type, and invalidation by either still reaches them. Invalidation by cache tag misses them until they expire.
  • A failed request skips everything that depends on the merged fetch. When the request for the cold aliases fails to reach the subgraph, the warm aliases keep their cached data and only the cold ones report the failure. Fetches that depend on the merged fetch are skipped as a whole, including those that depend only on a warm alias. Their fields stay null without an error of their own.

Cache tags

A subgraph names what its response is about by returning a cache tag extension. Tags are recorded as a secondary index over the cached entries, so an entry can be found by what it is about rather than only by the key it happens to sit under. The shape depends on what was fetched, because an entity fetch caches one entry per entity while a root query fetch caches its whole answer as one. Entity fetches return apolloEntityCacheTags, an array of arrays holding one list per entity, in the same order as the _entities answered:
Root query fetches return apolloCacheTags, one flat array:
An outer array whose length does not match the entities answered costs the whole response its tags, rather than being zipped as far as the shorter of the two. Either extension is consumed by the router and never forwarded to clients. The router does not record tags from a merged entity fetch. See Limitations.

Invalidation

Cached entries expire when their TTL ends. Send a request to the invalidation endpoint on its separate port to remove entries earlier.
An entry written while an invalidation runs can survive it. The index lists the entry before its value lands, so the request finds nothing to remove and leaves it in place. The entry then serves data fetched before the invalidation until its TTL ends. Send the request again once in-flight writes have settled. This is a known bug and a fix is in progress.

Enable the endpoint

The endpoint only starts when the response cache itself is enabled.
config.yaml
The endpoint speaks plain HTTP. Keep listen_addr on the loopback interface, or terminate TLS in front of it, so the shared key is not sent in the clear.

Send a request

POST a JSON array to the endpoint, with the shared key as the Authorization header. There is no Bearer prefix.
Each element names one of three kinds. A successful request answers 202 with the number of cache keys removed:
The whole array is validated before anything is removed. One bad element refuses the request with 400 and removes nothing, so it can be corrected and resent as a whole. Each error names the element it is about:

Restrictions

Tags can be outdated. With the Redis provider, an entry cached again under different tags stays listed under its earlier tags until it expires. Invalidating a tag the entry no longer carries still removes it. With the memory provider, an entry is listed under the tags of its latest write only. Everything is scoped to a subgraph. A type or a cache tag is indexed per subgraph that answered, so every request names one. There is no way to invalidate a type or a tag across all subgraphs in one element. List the subgraphs instead. The memory provider is per replica. Each router holds its own cache, so an invalidation reaches only the replica it was sent to. Do not use the in-memory cache with multiple router instances in production.

CDN purging

A CDN in front of the router can cache whole responses and purge them by tag. Enable cache_tag_header and the router returns the tags of each response in a header the CDN reads. Set the header name and delimiter the CDN expects. Cloudflare purges by Cache-Tag with tags separated by a comma, the default. Fastly purges by Surrogate-Key with tags separated by a space.
config.yaml
The header names what a response is about. It does not tell the CDN to cache it. The CDN caches on the Cache-Control the router returns, which the cache control policy sets. On a hit, the lifetime reported for the served fetch is the time left on the entry, so the CDN never holds a response longer than the router does.

What the header carries

Every cacheable fetch of the request contributes three tiers of tag, coarsest first: A root query fetch has no single type and contributes no type tag. Declared tags are passed through unchanged. Tags are unique within the header. The subgraph- and type- prefixes are not reserved. A declared tag such as subgraph-products is indistinguishable in the header from the generated one, and a purge by it removes both. Avoid those prefixes in declared tags so a header tag maps back to one invalidation element.

When it is sent

The header is sent on cache hits and misses alike. A hit rebuilds it from what was stored with the entry, so both carry the same value. A fetch that was not cacheable contributes nothing, and a response with no cacheable fetch has no header. The header is not sent on subscriptions, deferred or batched responses, or responses carrying subgraph errors. A response with subgraph errors is marked no-store instead. The header does not depend on the invalidation indexes. It is built even when every index is off.

Size

max_bytes caps the header value. When the tags do not fit, they are grouped coarsest first, subgraph then type then declared, and the finest that do not fit are left out. A purge by a coarse tag still reaches the response. The default of 16384 bytes is Cloudflare’s limit on Cache-Tag. A declared tag that contains the delimiter, or a byte not valid in a header, is left out. max_bytes caps the whole value, not one tag. Fastly reads at most 1,024 bytes per key and ignores a longer key and every key after it, even when the value fits. Keep each declared tag at or below 1,024 bytes when purging through Fastly.

Purge both caches

A purge at the CDN removes the response from the CDN only. The next request reaches the router, which may still serve the entry from its own cache. Send the same change to the invalidation endpoint as well. Each header tag maps to one kind: Invalidate the router first, then purge the CDN, so the CDN’s refetch does not pick up the stale entry.

Validation

Startup fails when name is not a valid header name or is Content-Length, Content-Type or Transfer-Encoding, when delimiter is empty or not valid in a header, or when max_bytes is not greater than zero.

Observability

The router logs once at startup when the cache is enabled. With the Redis provider:
With the memory provider:
fallback_ttl is the value under all and subgraph_overrides counts the entries under subgraphs. Nothing is logged when the cache is disabled. Cache failures never fail a request. When a read or a write fails, the router serves from the subgraph and logs a warning:
That warning is sampled to one line per second. An unreachable cache produces one failure per cacheable fetch of every request in flight, and logging it unsampled would compound the outage.
  • Storage Providers defines the Redis instance the redis provider references.
  • Adjusting Cache Control sets the Cache-Control the client and any CDN in front of the router receive.
  • Cache Warmer is a different cache. It pre-populates the operation plan cache and does not cache subgraph responses.