Experimental feature. Response caching is in alpha and is subject to change, and should not be used in production environments.
_entities root field to resolve fields of an entity that another subgraph owns. Entries are keyed per entity and per selection set, so asking for different fields of the same entity does not share an entry.
Root query fetches. A root query fetch is the request the router sends a subgraph for the root fields the client selected. The whole answer is stored as one entry. The key covers the operation text and every variable value. Any difference in either is a miss, and the request goes to the subgraph as it would have. Only queries are cached. Mutations and subscriptions are never cached.
Response caching is disabled by default.
What gets cached
A subgraph response is cached only when itsCache-Control header asks to be. The rules apply in this order:
no-storerefuses caching.no-cacherefuses caching, in any form. This includes the qualifiedno-cache="Set-Cookie"form.privatemarks the response as belonging to a private context. It is cached only under a key scoped to that context, and only whenprivate_idis configured and resolves for the request. See Private responses. Without one,privaterefuses caching.s-maxagesets the lifetime and takes precedence overmax-age.max-agesets the lifetime whens-maxageis absent.- When neither lifetime is present,
fallback_ttlapplies if the header containspublic,private,must-revalidate,proxy-revalidate,stale-if-error, orstale-while-revalidate. The stale directives require a valid duration, such asstale-if-error=60.
public is optional. For example, Cache-Control: max-age=60 and Cache-Control: s-maxage=60 both allow caching for 60 seconds.
A value of 0 for the selected lifetime refuses caching. It does not fall through to the next rule. For example, max-age=60, s-maxage=0 is not cached, while max-age=0, s-maxage=60 is cached for 60 seconds. An invalid duration in a recognized directive also prevents caching.
A missing or empty Cache-Control header is never cached. Headers containing only unrecognized directives, such as cdn-cache-control=60 or immutable, are also never cached. fallback_ttl does not apply to these responses.
Revalidation and stale-response directives opt into the fallback lifetime. The response cache does not revalidate entries or serve expired entries.
A response carrying GraphQL errors is never cached, whatever its Cache-Control header says.
Configuring the response cache does not make anything cacheable on its own. Subgraphs opt in with a recognized caching directive, such as
Cache-Control: max-age=60 or Cache-Control: public. If nothing appears to be cached, check the subgraph response headers first.Request directives
A client controls the cache for its own request through theCache-Control request header. Two directives are honored:
no-cacheis never answered from the cache. Every fetch of the operation goes to its subgraph, and the responses are stored as they are for any other request, so ano-cacherequest refreshes the entries it bypassed.no-storekeeps the responses to this request out of the cache. Nothing is stored, but a warm entry still answers the request, as RFC 9111 allows.
no-cache, no-store neither reads nor writes the cache. Every other request directive is ignored, including max-age, min-fresh, max-stale and only-if-cached, as is the Pragma header. A Cache-Control request header that does not parse is ignored as a whole.
Enable it
Redis
Every router replica shares one cache and entries outlive the process. Define a storage provider, then reference it byprovider_id.
Redis 7.0 or newer is required to use this feature.
provider_id fails startup.
Router memory
Each replica holds its own cache, nothing survives a restart, and no Redis is needed.provider defaults to redis, so caching in memory has to be asked for by name. This prevents a missing provider_id from quietly turning one shared cache into one cache per replica.
Configuration reference
config.yaml
storage is required when enabled is true.
Environment variables bypass the config schema validation. In each enabled scope, a fallback_ttl of zero or less fails startup, as does a private_id that does not compile.
Per-subgraph configuration
all applies to every subgraph. An entry under subgraphs, keyed by subgraph name, replaces all for that subgraph:
config.yaml
inventory is never cached, products keeps responses without a lifetime for five minutes, and every other subgraph keeps them for thirty seconds.
An entry is a full override and is explicit. Nothing under all reaches a subgraph that has an entry, and no key is defaulted:
enabledmust be set. An entry without it does not cache that subgraph. Afallback_ttlon its own does not enable anything.- An enabled entry must set
fallback_ttl, or startup fails naming the entry. - An entry without
private_idhas none. Soproductsabove does not cache itsprivateresponses, althoughallnames an id. Repeat theprivate_idin the entry to keep it.
all.enabled to false turns the cache off for every subgraph without an entry, so only the entries that set enabled: true are cached. With no such entry the router starts and logs a warning, and nothing is cached.
An entry naming a subgraph the graph does not have is kept and logged at startup, since a feature flag graph may have it.
The invalidation indexes, the invalidation endpoint, and the cache tag header are not per subgraph.
Private responses
A subgraph whose answer depends on who is asking marks itCache-Control: private. A shared cache must not serve that answer to anyone else, but the router can still cache it for the context it was produced in, so the same context asking again is served without a subgraph request.
To do that the router needs an id for that context. private_id is an expression evaluated once per request that yields it, such as a JWT claim or a request header. The id partitions the cache key: what it names is up to the expression, and it need not be a single user. A tenant, an organisation, or a role are all valid contexts, as long as every request that shares the id may share the response.
config.yaml
all it applies to every subgraph. Under a subgraphs entry it applies to that subgraph alone, and an entry without one has none. See Per-subgraph configuration.
With a private_id in place:
- A
privateresponse is stored under a key that includes the id, so each context has an entry of its own. A lookup with one id never finds an entry stored under another. - A request for which the expression yields nothing, such as an anonymous request whose claim is
nil, is not cached when the subgraph answersprivate. The subgraph is asked every time. This is not an error. publicresponses are unaffected. They are stored under the shared key and served to every request, whatever the id.- With
cache_control_policyor amost_restrictive_cache_controlresponse header rule in place, a response served from a private entry reaches the client withCache-Control: privateand the entry’s remaining lifetime, so a browser or CDN in front of the router does not share it either. Without one, the router sends noCache-Controlfor the hit. - Invalidation by subgraph, type, or cache tag drops private entries like any other.
Varying responses
A subgraph whose answer depends on a request header, such asAccept-Language, names it in a Vary response header. The router then keeps one entry per set of values it sent, learned from the first response, the way an HTTP proxy cache does:
- The first request for an operation misses. The names in the response’s
Varyheader are stored at the entry’s key as a record, and the body is stored under a variant key derived from the values the router sent. - A later lookup finds the record, digests the values this request sends the subgraph for those headers, and looks the variant up under the resulting key. A varying response costs two cache round trips per lookup. A response without
Varycosts one, as before. - Values are taken from the request as the subgraph receives it, after header propagation. A header named in
Varythat the router does not forward counts as empty, so it never splits the cache. Vary: *is never cached, and neither is a response naming more than 100 headers.- A header whose value differs per caller, such as
CookieorAuthorization, yields one variant per caller. Nobody is served another caller’s body, but the entries are rarely reused. A subgraph whose answer is per caller should sayprivateinstead. - A subgraph that changes what it varies on adds the new set of names to the record on its next miss, newest first, rather than replacing it. A lookup tries one variant per set in that order and serves the first it finds, so variants stored under the old names stay reachable until they expire. A record keeps at most 8 sets; past that the oldest are dropped.
- With
engine.enable_multi_fetchon, a merged entity fetch carries oneVaryfor the whole response, and the router applies it to every alias in it. An alias whose own answer would have varied on less is stored under the wider set, which splits its entries by headers they do not depend on. That costs extra misses, never a wrong body. A merged fetch follows records the same way a single fetch does, with one second round trip for all of its aliases together. The lever for the cost is on the subgraph side: emit the minimumVaryfor what was actually resolved rather than a fixed union for the service.
Limitations
These apply to the current alpha. A batch is served from the cache only when every entity in it is present. One miss sends the whole batch to the subgraph, including the entities that were already cached. Entities are stored individually, so a later batch made up only of cached entities is a full hit. Forwarded request headers reach the cache key only throughVary. The key covers the request the router renders for the subgraph, which is the operation text and its variables. Headers added by header propagation are applied after that and never reach the key on their own. Two requests that differ only by a propagated header share one entry unless the subgraph names that header in Vary, as described under Varying responses. A subgraph that varies by request header without saying so must use no-store, or private together with a private_id whose value accounts for that header. Omitting public does not prevent caching.
Cache hits produce no subgraph telemetry. A hit skips the subgraph fetch, so no subgraph span or metric is recorded for it. Subgraph request counts fall as the hit rate rises. Use them to confirm the cache is working, not to measure traffic.
Merged entity fetches have their own rules. With engine.enable_multi_fetch on, entity fetches to the same subgraph within one parallel wave are merged into a single request with aliased _entities fields. The response is taken apart again on the way into the cache, so each alias stores one entry per entity, the same as the unmerged fetch it replaces would have. From there:
- Each alias is served all-or-nothing, like a batch. A warm alias is left out of the next merged request. A cold one is fetched whole.
- Only a request that finds every alias warm is a cache hit. When some aliases are warm and others are not, the router still sends the request for the cold ones.
- The reported
Cache-Controlis capped at the warm aliases. The lifetime reported for a partly warm fetch is the life left on its warm aliases. - Declared cache tags are dropped.
apolloEntityCacheTagsis one flat list with nothing to say which alias each entry belongs to. The entries are still indexed by subgraph and type, and invalidation by either still reaches them. Invalidation by cache tag misses them until they expire. - A failed request skips everything that depends on the merged fetch. When the request for the cold aliases fails to reach the subgraph, the warm aliases keep their cached data and only the cold ones report the failure. Fetches that depend on the merged fetch are skipped as a whole, including those that depend only on a warm alias. Their fields stay
nullwithout an error of their own.
Cache tags
A subgraph names what its response is about by returning a cache tag extension. Tags are recorded as a secondary index over the cached entries, so an entry can be found by what it is about rather than only by the key it happens to sit under. The shape depends on what was fetched, because an entity fetch caches one entry per entity while a root query fetch caches its whole answer as one. Entity fetches returnapolloEntityCacheTags, an array of arrays holding one list per entity, in the same order as the _entities answered:
apolloCacheTags, one flat array:
Invalidation
Cached entries expire when their TTL ends. Send a request to the invalidation endpoint on its separate port to remove entries earlier.An entry written while an invalidation runs can survive it. The index lists the entry before its value lands, so the request finds nothing to remove and leaves it in place. The entry then serves data fetched before the invalidation until its TTL ends. Send the request again once in-flight writes have settled. This is a known bug and a fix is in progress.
Enable the endpoint
The endpoint only starts when the response cache itself is enabled.config.yaml
listen_addr on the loopback interface, or terminate TLS in front of it, so the shared key is not sent in the clear.
Send a request
POST a JSON array to the endpoint, with the shared key as the Authorization header. There is no Bearer prefix.
A successful request answers
202 with the number of cache keys removed:
400 and removes nothing, so it can be corrected and resent as a whole. Each error names the element it is about:
Restrictions
Tags can be outdated. With the Redis provider, an entry cached again under different tags stays listed under its earlier tags until it expires. Invalidating a tag the entry no longer carries still removes it. With the memory provider, an entry is listed under the tags of its latest write only. Everything is scoped to a subgraph. A type or a cache tag is indexed per subgraph that answered, so every request names one. There is no way to invalidate a type or a tag across all subgraphs in one element. List the subgraphs instead. The memory provider is per replica. Each router holds its own cache, so an invalidation reaches only the replica it was sent to. Do not use the in-memory cache with multiple router instances in production.CDN purging
A CDN in front of the router can cache whole responses and purge them by tag. Enablecache_tag_header and the router returns the tags of each response in a header the CDN reads. Set the header name and delimiter the CDN expects. Cloudflare purges by Cache-Tag with tags separated by a comma, the default. Fastly purges by Surrogate-Key with tags separated by a space.
config.yaml
The header names what a response is about. It does not tell the CDN to cache it. The CDN caches on the
Cache-Control the router returns, which the cache control policy sets. On a hit, the lifetime reported for the served fetch is the time left on the entry, so the CDN never holds a response longer than the router does.What the header carries
Every cacheable fetch of the request contributes three tiers of tag, coarsest first:
A root query fetch has no single type and contributes no type tag. Declared tags are passed through unchanged. Tags are unique within the header.
The
subgraph- and type- prefixes are not reserved. A declared tag such as subgraph-products is indistinguishable in the header from the generated one, and a purge by it removes both. Avoid those prefixes in declared tags so a header tag maps back to one invalidation element.
When it is sent
The header is sent on cache hits and misses alike. A hit rebuilds it from what was stored with the entry, so both carry the same value. A fetch that was not cacheable contributes nothing, and a response with no cacheable fetch has no header. The header is not sent on subscriptions, deferred or batched responses, or responses carrying subgraph errors. A response with subgraph errors is markedno-store instead.
The header does not depend on the invalidation indexes. It is built even when every index is off.
Size
max_bytes caps the header value. When the tags do not fit, they are grouped coarsest first, subgraph then type then declared, and the finest that do not fit are left out. A purge by a coarse tag still reaches the response. The default of 16384 bytes is Cloudflare’s limit on Cache-Tag.
A declared tag that contains the delimiter, or a byte not valid in a header, is left out.
max_bytes caps the whole value, not one tag. Fastly reads at most 1,024 bytes per key and ignores a longer key and every key after it, even when the value fits. Keep each declared tag at or below 1,024 bytes when purging through Fastly.
Purge both caches
A purge at the CDN removes the response from the CDN only. The next request reaches the router, which may still serve the entry from its own cache. Send the same change to the invalidation endpoint as well. Each header tag maps to onekind:
Invalidate the router first, then purge the CDN, so the CDN’s refetch does not pick up the stale entry.
Validation
Startup fails whenname is not a valid header name or is Content-Length, Content-Type or Transfer-Encoding, when delimiter is empty or not valid in a header, or when max_bytes is not greater than zero.
Observability
The router logs once at startup when the cache is enabled. With the Redis provider:fallback_ttl is the value under all and subgraph_overrides counts the entries under subgraphs.
Nothing is logged when the cache is disabled.
Cache failures never fail a request. When a read or a write fails, the router serves from the subgraph and logs a warning:
Related
- Storage Providers defines the Redis instance the
redisprovider references. - Adjusting Cache Control sets the
Cache-Controlthe client and any CDN in front of the router receive. - Cache Warmer is a different cache. It pre-populates the operation plan cache and does not cache subgraph responses.