Files
ResolveSpec/audit/pkg/cache.audit.md
T
Hein bc8bff7955
Tests / Unit Tests (push) Failing after 24s
Tests / Integration Tests (push) Failing after 26s
Build , Vet Test, and Lint / Build (push) Successful in 1m7s
Build , Vet Test, and Lint / Run Vet Tests (1.23.x) (push) Successful in 1m31s
Build , Vet Test, and Lint / Lint Code (push) Successful in 1m33s
Build , Vet Test, and Lint / Run Vet Tests (1.24.x) (push) Successful in 1m34s
docs(audit): add audit reports for pkg/testmodels and pkg/tracing
2026-09-29 17:15:00 +02:00

1360 lines
56 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Audit: `pkg/cache`
| | |
|---|---|
| **Package** | `github.com/bitechdev/ResolveSpec/pkg/cache` |
| **Files** | `cache.go` (76), `cache_manager.go` (167), `provider.go` (65), `provider_memory.go` (342), `provider_memcache.go` (284), `provider_redis.go` (269), `example_usage.go` (266), `cache_test.go` (69) |
| **Audit date** | 2026-09-29 |
| **Axes** | thread locking/waiting, slowness, security, panic handling & logging |
| **Threat model** | hostile internet client; request bodies, headers, query params, schema/table/column names all attacker-controlled |
| **Depth** | deep (hot package) |
## Summary
`pkg/cache` is a thin `Provider` abstraction over three backends. Two things
make it much more security-relevant than a cache normally is:
1. **It stores authentication state.** `pkg/security/providers.go:398-402` caches
`UserContext` under the key `fmt.Sprintf("auth:session:%s", token)` — the raw
bearer token from the `Authorization` header. So **cache keys are directly
attacker-controlled**, the cached value is an authorization decision, and
`DeleteByPattern` is the only session-revocation mechanism.
2. **It is never explicitly initialized.** Nothing in `pkg/` calls
`Initialize`/`UseMemory`/`UseRedis`/`UseMemcache`; every consumer reaches the
cache through `GetDefaultCache()` (`cache.go:48`), which lazily constructs a
`MemoryProvider` **on the request path** without synchronization.
The most serious findings are: a cache-write failure being converted into an
authentication failure (`GetOrSet`, `cache_manager.go:126`), session revocation
that silently cannot work on memcache, an unsynchronized lazy singleton read by
concurrent HTTP handlers, an unbounded `tagToKeys` index that `MaxSize` does not
cover, and `MemoryProvider.Get` taking the **write** lock on every cache hit.
There is **no `recover()` anywhere in this package** and no logging at all —
grep for `recover()` and `logger.` across all 1 538 lines returns zero hits.
Every error is either returned to the caller or discarded with `_ =`.
## Findings
| # | Severity | Axis | Finding |
|---|---|---|---|
| 1 | **Critical** | security | Cache-write failure is returned as an error from `GetOrSet`, so a cache outage becomes a total authentication outage |
| 2 | **Critical** | security | Session revocation (`DeleteByPattern`) is unimplemented on memcache and returns an error — revoked sessions keep authenticating |
| 3 | **High** | locking | `defaultCache` is an unsynchronized lazy singleton read/written from concurrent request handlers |
| 4 | **High** | security / slowness | `tagToKeys` is unbounded and leaked by six paths; `MaxSize` bounds only `items` |
| 5 | **High** | locking / slowness | `MemoryProvider.Get` acquires the **write** lock on every hit |
| 6 | **High** | security | `DeleteByPattern` pattern syntax differs per provider (Go regexp vs Redis glob vs error); memory matches unanchored |
| 7 | **High** | security | Raw session token embedded in the `"key not found: %s"` error string |
| 8 | **High** | security | No single-flight: concurrent misses on one key run N loaders (auth stampede against the session stored procedure) |
| 9 | **Medium** | security | Attacker-controlled keys break memcache's 250-byte/no-whitespace key rule; error is swallowed as a miss, then fails the write |
| 10 | **Medium** | correctness | TOCTOU in `MemoryProvider.Get`: expired-item path deletes unconditionally after dropping the read lock |
| 11 | **Medium** | correctness | `MemoryProvider.Close()` sets `items = nil`; a subsequent `Set` panics and nothing recovers |
| 12 | **Medium** | security | Memcache tag index is non-atomic read-modify-write and shares the key namespace; lost updates silently drop keys from invalidation |
| 13 | **Medium** | security | `Clear()` maps to `FlushAll()` / `FlushDB()` — wipes the whole shared server/DB |
| 14 | **Medium** | correctness | `MemcacheProvider.Close()` is a no-op justified by a false comment; gomemcache **does** have `Close()` |
| 15 | **Medium** | security | Redis and memcache have no TLS option at all; `AUTH` password and cached `UserContext` cross the wire in cleartext |
| 16 | **Medium** | correctness | `Cache.Remember` returns a different Go type on hit vs miss |
| 17 | **Medium** | slowness | `Get`/`Exists` swallow **all** backend errors as a cache miss — a degraded backend is invisible and stampedes the DB |
| 18 | **Low** | slowness | `evictOne` is an O(n) scan under the write lock, run per insertion at capacity |
| 19 | **Low** | correctness | `CleanExpired` is dead code; there is no janitor, so expired items are only reclaimed on access |
| 20 | **Low** | security | `RedisProvider.Stats` returns the raw `INFO` output in `ProviderStats["info"]` |
| 21 | **Low** | correctness | `ctx` is accepted and completely ignored by the memcache provider |
| 22 | **Low** | correctness | `MaxSize <= 0` disables eviction entirely — an unbounded in-memory cache |
| 23 | **Low** | correctness | Provider constructors mutate the caller's config struct |
| 24 | **Low** | correctness | No defensive copy of `[]byte` on `Set`/`Get` in the memory provider |
| 25 | **Low** | hygiene | `example_usage.go` ships `log.Fatal` calls in a library package |
---
### 1. Critical — a cache-write failure is an authentication failure
`pkg/cache/cache_manager.go:112-141`:
```go
func (c *Cache) GetOrSet(ctx context.Context, key string, dest interface{}, ttl time.Duration, loader func() (interface{}, error)) error {
err := c.Get(ctx, key, dest)
if err == nil {
return nil
}
value, err := loader()
if err != nil {
return fmt.Errorf("loader failed: %w", err)
}
// Store in cache
if err := c.Set(ctx, key, value, ttl); err != nil {
return fmt.Errorf("failed to cache value: %w", err) // <-- line 127
}
...
}
```
The loader has already succeeded — the authoritative value is in hand — but a
failure to *cache* it aborts the whole call. The only consumer of `GetOrSet` is
authentication, `pkg/security/providers.go:398-444`:
```go
cacheKey := fmt.Sprintf("auth:session:%s", token)
var userCtx UserContext
err := a.cache.GetOrSet(r.Context(), cacheKey, &userCtx, a.cacheTTL, func() (any, error) {
// ... queries the session stored procedure, returns &user on success
})
if err != nil {
lastErr = err
continue // Try next token
}
```
Any error — including "failed to cache value" — is treated as *this token is not
valid*, and after the token loop the request is rejected.
**Failure scenario.** Redis is configured and becomes unreachable (restart,
failover, network partition, `maxmemory` reached with `noeviction`).
`RedisProvider.Get` swallows the error and reports a miss
(`provider_redis.go:91-93`), the loader runs and the database confirms the
session is valid, then `RedisProvider.Set` returns the connection error and
`GetOrSet` returns it. **Every request from every user is now rejected with an
authentication error**, even though both the database and the sessions are
healthy. A cache is supposed to be a latency optimization; here it is a hard
dependency of the auth path, and its failure mode is total outage. The same
applies to `maxmemory` pressure, which an attacker can induce (see finding 4).
**Recommendation.** A cache-write failure must be non-fatal. Log it and return
the loaded value:
```go
if err := c.Set(ctx, key, value, ttl); err != nil {
logger.Warn("cache: failed to store key (continuing uncached): %v", ctx, err)
}
```
Separately, `pkg/security` should not conflate "cache layer failed" with
"credential rejected"; the loader's own error is the only one that should fail
authentication. Consider having `GetOrSet` return the loaded value plus a
non-fatal cache error, or wrap cache errors in a sentinel the caller can test
with `errors.Is`.
---
### 2. Critical — session revocation silently cannot work on memcache
`pkg/security/providers.go:463-479` is the only session-revocation path:
```go
func (a *DatabaseAuthenticator) ClearCache(token string) error {
ctx := context.Background()
if token != "" {
cacheKey := fmt.Sprintf("auth:session:%s", token)
return a.cache.Delete(ctx, cacheKey)
}
// Clear all auth cache entries
return a.cache.DeleteByPattern(ctx, "auth:session:*")
}
func (a *DatabaseAuthenticator) ClearUserCache(userID int) error {
ctx := context.Background()
pattern := "auth:session:*"
return a.cache.DeleteByPattern(ctx, pattern)
}
```
`pkg/cache/provider_memcache.go:249-254`:
```go
// DeleteByPattern removes all keys matching the pattern.
// Note: Memcache does not support pattern-based deletion natively.
// This is a no-op for memcache and returns an error.
func (m *MemcacheProvider) DeleteByPattern(ctx context.Context, pattern string) error {
return fmt.Errorf("pattern-based deletion is not supported by Memcache")
}
```
**Failure scenario.** A deployment uses memcache (`cache.provider: memcache`).
An account is compromised; an operator disables the user or the sessions are
revoked in the database, and the application calls `ClearUserCache(userID)`.
That returns an error and **removes nothing**. The attacker's cached
`UserContext` continues to authenticate every request for the full
`a.cacheTTL` — the database is never consulted again during that window
(`GetOrSet` short-circuits on a cache hit). Whether the operator even learns
this failed depends entirely on whether the caller checks the returned error;
`ClearUserCache` is also broken for a second reason — it ignores `userID` and
would have revoked every session in the process.
Note also that `ClearUserCache`'s pattern is *not* user-scoped, so even on Redis
and memory it is a global logout, not a per-user one. That direction is at least
fail-safe.
**Recommendation.** Revocation must not depend on a capability the provider may
not have. Options, in order of preference:
- Tag every session entry (`SetWithTags` with tags `auth:session`,
`auth:user:<id>`) and revoke with `DeleteByTag`, which all three providers
implement. This also makes `ClearUserCache` actually per-user.
- Keep a short `cacheTTL` (seconds, not minutes) so the revocation window is
bounded regardless.
- Make `DeleteByPattern`'s unsupported case loud: have `pkg/security` refuse to
start, or fall back to `Clear`, when the configured provider cannot revoke.
- At minimum, log at error level when a revocation call fails.
---
### 3. High — `defaultCache` is an unsynchronized lazy singleton on the request path
`pkg/cache/cache.go:9-62`:
```go
var (
defaultCache *Cache
)
func Initialize(provider Provider) {
defaultCache = NewCache(provider)
}
func UseMemory(opts *Options) error {
provider := NewMemoryProvider(opts)
defaultCache = NewCache(provider)
return nil
}
// ... UseRedis (:32), UseMemcache (:42) likewise
func GetDefaultCache() *Cache {
if defaultCache == nil {
_ = UseMemory(&Options{
DefaultTTL: 5 * time.Minute,
MaxSize: 10000,
})
}
return defaultCache
}
func SetDefaultCache(cache *Cache) {
defaultCache = cache
}
```
Six functions write `defaultCache` and `GetDefaultCache` both reads and writes
it, with no mutex, no `sync.Once` and no `atomic.Pointer`. `GetDefaultCache` is
called **from HTTP request handlers**: `pkg/restheadspec/handler.go:833`
(`cache.GetDefaultCache().Get(ctx, cacheKey, cachedTotalData)`),
`pkg/restheadspec/cache_helpers.go:109` and `:118`, and the equivalent
`pkg/resolvespec` paths.
Nothing in `pkg/` ever calls `Initialize` or `Use*`, so in a default deployment
**the first traffic to arrive is what initializes the cache**, concurrently.
**Failure scenario.** Two requests arrive simultaneously on a cold process.
Both observe `defaultCache == nil`, both run `UseMemory`, each constructing its
own `MemoryProvider`. One assignment wins. Request A writes its query total into
the provider that loses and is immediately garbage — so the entry is
unreachable, the cache reports a permanent miss for it, and the count is
recomputed from the database on every subsequent request. Worse, if this races
with an application's explicit `UseRedis` during startup, the Redis provider can
be clobbered by the lazy memory provider (or vice versa) and **the process
silently runs on the wrong backend**, which for `pkg/security` means session
cache entries that no other process shares and that `ClearCache` on another
instance can never reach.
This is also a genuine data race on the pointer word: unsynchronized
read/write of `defaultCache`, which `go test -race` would report immediately.
See `_CROSS-CUTTING.audit.md` — `-race` is never run in this repo, and
`pkg/cache` is not in the tested package set.
Note the same pattern exists in `pkg/security/providers.go:145`
(`cacheInstance = cache.GetDefaultCache()`), which at least happens at
construction time.
**Recommendation.** Guard the global with `sync.RWMutex` or store it in an
`atomic.Pointer[Cache]`, and make the lazy default a `sync.Once`:
```go
var (
defaultCache atomic.Pointer[Cache]
defaultOnce sync.Once
)
func GetDefaultCache() *Cache {
if c := defaultCache.Load(); c != nil {
return c
}
defaultOnce.Do(func() {
defaultCache.CompareAndSwap(nil, NewCache(NewMemoryProvider(&Options{
DefaultTTL: 5 * time.Minute, MaxSize: 10000,
})))
})
return defaultCache.Load()
}
```
Also: every replacement path drops the previous provider **without closing it**,
so `UseRedis` after a lazy `UseMemory` (or two `UseRedis` calls) leaks the old
provider's connection pool and, for Redis, its background goroutines. Close the
old provider on swap.
---
### 4. High — `tagToKeys` is unbounded; `MaxSize` bounds only `items`
`MemoryProvider` holds two maps (`provider_memory.go:30-37`):
```go
type MemoryProvider struct {
mu sync.RWMutex
items map[string]*memoryItem
tagToKeys map[string]map[string]struct{} // tag -> set of keys
options *Options
...
}
```
`MaxSize` is checked only against `len(m.items)` (`:105`, `:135`). Six paths
remove entries from `items` **without** removing them from `tagToKeys`:
| Path | Line | Cleans `tagToKeys`? |
|---|---|---|
| `Get` — expired-item delete | `:70` | no |
| `Set` — overwrites a tagged key | `:111` | no (and drops `Tags`, so the entry becomes unreachable for cleanup) |
| `evictOne` — expired scan | `:313` | no |
| `evictOne` — LRU victim | `:324` | no |
| `DeleteByPattern` | `:242` | no |
| `Clear` | `:254` | no — `m.tagToKeys` is never reset |
| `CleanExpired` | `:336` | no |
Only `Delete` (`:178-187`) and `SetWithTags` (`:141-151`) maintain it, and
`DeleteByTag` (`:226`) drops one whole tag.
Tags come from `pkg/restheadspec/cache_helpers.go:99-105`:
```go
func buildCacheTags(schema, tableName string) []string {
return []string{
fmt.Sprintf("schema:%s", strings.ToLower(schema)),
fmt.Sprintf("table:%s", strings.ToLower(tableName)),
}
}
```
and keys from `buildExtendedQueryCacheKey` (`:43-85`) — a SHA-256 of the full
query shape, including filters, sort, `customWhere`, `customOr`, `customJoin`,
expand and cursors.
**Failure scenario.** An attacker issues `GET /api/public/orders?...` in a loop,
varying one filter value each time. Every request produces a distinct SHA-256
key, `setQueryTotalCache` stores it under the tags `schema:public` and
`table:orders`, and `tagToKeys["table:orders"][key]` gains a member. Once
`items` reaches `MaxSize` (10 000 by default), `evictOne` starts discarding
items — but **never** the corresponding `tagToKeys` members. `items` stays
capped at 10 000; `tagToKeys["table:orders"]` grows by one 64-character key per
request, forever. At roughly 100 bytes per map entry, a few million requests —
easily reachable at modest rate — costs hundreds of megabytes of heap that
nothing will ever reclaim, because `Clear()` does not reset the map and
`CleanExpired` does not touch it. This is a **memory-exhaustion DoS driven
purely by query-string variation**, and it is cheap for the attacker: the
expensive part (the actual count query) can be avoided by hitting a table whose
count is trivial.
The leak also breaks invalidation correctness: `DeleteByTag` iterates a key set
full of keys that no longer exist, and a plain `Set` over a previously tagged
key leaves that key in the tag index while clearing its `Tags` — so the item can
be deleted by a tag it no longer claims.
**Recommendation.** Factor tag maintenance into a single private helper and call
it from every removal path:
```go
// caller must hold m.mu for writing
func (m *MemoryProvider) removeLocked(key string) {
if item, ok := m.items[key]; ok {
for _, tag := range item.Tags {
if ks := m.tagToKeys[tag]; ks != nil {
delete(ks, key)
if len(ks) == 0 {
delete(m.tagToKeys, tag)
}
}
}
}
delete(m.items, key)
}
```
Use it in `Get`'s expired path, `Set` (before overwrite), `evictOne`,
`DeleteByPattern` and `CleanExpired`; reset `m.tagToKeys` in `Clear`; and
account `len(m.tagToKeys)` (or total members) against an explicit bound.
Independently, cap the number of distinct tags and the members per tag.
---
### 5. High — `MemoryProvider.Get` takes the write lock on every hit
`provider_memory.go:56-88`:
```go
func (m *MemoryProvider) Get(ctx context.Context, key string) ([]byte, bool) {
// First try with read lock for fast path
m.mu.RLock()
item, exists := m.items[key]
...
value := item.Value
m.mu.RUnlock()
// Update access tracking with write lock
m.mu.Lock()
item.LastAccess = time.Now()
item.HitCount++
m.mu.Unlock()
m.hits.Add(1)
return value, true
}
```
The comment promises a read-lock fast path, but every **successful** lookup ends
in an exclusive lock to bump two bookkeeping fields. The `RWMutex` therefore
provides no read concurrency at all on the hot path, and Go's `RWMutex` blocks
*new* readers once a writer is waiting — so a burst of concurrent hits
degenerates into a fully serialized queue with two lock handoffs per operation.
**Failure scenario.** The session cache is the default `MemoryProvider`. Under
concurrent load, every authenticated request performs
`GetOrSet` → `Get` → cache hit → exclusive lock. With hundreds of in-flight
requests the mutex becomes the throughput ceiling for the entire API, and
because the write lock is taken *after* the read lock is released, each hit pays
two full lock acquisitions. Adding `HitCount` to a per-item `atomic.Int64`
would make this free; as written, the cache that exists to reduce latency is the
serialization point.
A secondary defect: between `RUnlock` at `:78` and `Lock` at `:81` the item may
have been deleted or replaced, so the code can mutate an orphaned struct. Harmless
but confirms the bookkeeping does not need the lock.
**Recommendation.** Make the counters lock-free and drop the write lock:
```go
type memoryItem struct {
Value []byte
Expiration time.Time
lastAccess atomic.Int64 // unix nanos
hitCount atomic.Int64
Tags []string
}
```
Then the whole `Get` runs under `RLock`. If exact LRU ordering matters, consider
an approximate clock (update `lastAccess` only if it is more than a second
stale) or a sharded map to cut contention.
---
### 6. High — `DeleteByPattern` has three incompatible pattern languages
The interface (`provider.go:29-31`) says only "Pattern syntax depends on the
provider implementation", and the three implementations diverge completely:
| Provider | Line | Semantics |
|---|---|---|
| memory | `provider_memory.go:235-244` | `regexp.Compile` + **unanchored** `MatchString` |
| redis | `provider_redis.go:197` | `SCAN MATCH` — Redis glob |
| memcache | `provider_memcache.go:252-254` | always an error |
```go
// memory
re, err := regexp.Compile(pattern)
if err != nil {
return fmt.Errorf("invalid pattern: %w", err)
}
for key := range m.items {
if re.MatchString(key) {
delete(m.items, key)
}
}
```
The one caller, `pkg/security/providers.go:470`, passes `"auth:session:*"` —
a Redis glob. Interpreted as a Go regexp that is `auth:session` followed by zero
or more `:`, matched unanchored, so it happens to match the intended keys (and
any key merely *containing* `auth:session`). It works by coincidence, not design.
**Failure scenario.** Two ways this bites:
- **Over-deletion.** Because matching is unanchored, a glob like `user:*` becomes
the regexp `user:*` = `user` + zero-or-more colons, which matches *any* key
containing `user` — including `auth:session:<token>` if a token happens to
contain that substring. Conversely a caller who writes a glob such as
`*` gets `regexp.Compile("*")` → `error parsing regexp: missing argument to
repetition operator`, i.e. a silent no-op invalidation where the author
expected a full flush. Stale authorization data continues to be served.
- **Attacker-influenced regexp.** `regexp.Compile` runs on the caller's string
**while holding the write lock** (`:232-238`), so if any future caller derives
a pattern from request input, an attacker both controls the compiled program
and blocks every other cache operation for its duration. Go's RE2 has no
catastrophic backtracking, but compilation of a large pattern is not free and
the lock is held across it.
Redis's side has its own cost: `r.client.Scan(ctx, 0, pattern, 0)` with
`count = 0` leaves the server at its default `COUNT 10`, so revoking sessions
walks the entire keyspace in ~10-key increments — thousands of round trips on a
large DB, executed synchronously inside `ClearCache`.
**Recommendation.** Define one pattern language in the interface — a glob is the
right choice, since it is the one Redis supports natively — and implement it for
memory with `path.Match` (or an explicit anchored translation to regexp),
compiled **before** taking the lock. Reject patterns the provider cannot honour
with a typed `ErrUnsupported` so callers can branch. Pass a sensible `COUNT`
(e.g. 500) to `SCAN`. Better still, replace pattern deletion with tag deletion
at the one call site (see finding 2).
---
### 7. High — the session token is embedded in an error string
`cache_manager.go:22-43`:
```go
func (c *Cache) Get(ctx context.Context, key string, dest interface{}) error {
data, exists := c.provider.Get(ctx, key)
if !exists {
return fmt.Errorf("key not found: %s", key)
}
...
}
func (c *Cache) GetBytes(ctx context.Context, key string) ([]byte, error) {
data, exists := c.provider.Get(ctx, key)
if !exists {
return nil, fmt.Errorf("key not found: %s", key)
}
return data, nil
}
```
The key for the session cache is `"auth:session:" + token` — the raw bearer
credential. Every cache miss therefore allocates an error whose text contains a
live secret. `GetOrSet` discards it, but this is a public API on a `*Cache` that
`pkg/security` holds directly, and `pkg/security/keystore_database.go:232` calls
`ks.cache.Get` on an API-key cache too.
**Failure scenario.** Any caller that does `logger.Error("cache lookup failed: %v", err)`
publishes the bearer token to the application log **and**, per
`audit/pkg/logger.audit.md` finding 2, forwards it verbatim to Sentry, where it
is retained by a third party with no scrubbing (`pkg/errortracking` has no
`BeforeSend` hook). A token in a log aggregator is a replayable credential for
the whole `cacheTTL` — longer, if the log outlives the session. Note that this
requires only one careless `%v` at a call site; the package is handing out the
loaded weapon.
**Recommendation.** Never interpolate cache keys into errors. Use a package
sentinel and let the caller decide what is safe to log:
```go
var ErrNotFound = errors.New("cache: key not found")
...
if !exists {
return ErrNotFound
}
```
Callers then use `errors.Is(err, cache.ErrNotFound)` instead of matching on
strings, which also fixes the miss/error conflation noted in finding 17. If a
key must appear in diagnostics, log a truncated hash of it. Separately,
`pkg/security` should key the cache on a SHA-256 of the token rather than the
token itself, exactly as `keystore_database.go` already does for API keys
(`keystoreCacheKey(hash)`, `:287`).
---
### 8. High — no single-flight: concurrent misses run N loaders
`GetOrSet` (`cache_manager.go:112`) and `Remember` (`:145`) both do
check → load → store with nothing serializing concurrent callers on the same
key.
**Failure scenario (auth).** A client opens 200 connections with the same fresh
session token. All 200 miss the cache, all 200 enter the loader, and all 200
execute the session stored procedure
(`SELECT p_success, p_error, p_user::text FROM <session_fn>($1, $2)`,
`pkg/security/providers.go:413-415`) concurrently. Each success then also spawns
`go a.updateSessionActivity(...)` (`:447`), i.e. 200 more goroutines each issuing
a database write. The cache provides no protection at all for the first
round-trip, and under sustained concurrency where request arrival outpaces query
latency it never catches up: throughput is bounded by the database, not the
cache. A single valid credential is enough to drive this — no privilege needed.
**Failure scenario (query totals).** The same shape applies to
`restheadspec/handler.go:833`: a burst of identical expensive `COUNT(*)` queries
all miss together and all hit the database.
This is the classic cache stampede / thundering herd, and it is the reason
single-flight exists.
**Recommendation.** Wrap the loader in `golang.org/x/sync/singleflight`, which is
already an indirect dependency of most Go service stacks:
```go
type Cache struct {
provider Provider
sf singleflight.Group
}
func (c *Cache) GetOrSet(ctx context.Context, key string, dest any, ttl time.Duration, loader func() (any, error)) error {
if err := c.Get(ctx, key, dest); err == nil {
return nil
}
v, err, _ := c.sf.Do(key, func() (any, error) {
// re-check under the flight, then load and store
...
})
...
}
```
Note the group must be keyed per-`Cache`, and `Forget` should be called on
loader error so a failure is not shared beyond the in-flight set. For the auth
path specifically, also bound `updateSessionActivity` — an unbounded `go` per
request is its own DoS vector (raised again in the `pkg/security` audit).
---
### 9. Medium — attacker-controlled keys violate memcache's key rules
gomemcache enforces the protocol limits (verified in
`gomemcache@v0.0.0-20260422231931-4d751bb6e37c/memcache.go:58-91`):
```go
// ErrMalformedKey is returned when an invalid key is used.
// Keys must be at maximum 250 bytes long and not
// contain whitespace or control characters.
ErrMalformedKey = errors.New("malformed: key is too long or contains invalid characters")
func legalKey(key string) bool {
if len(key) > 250 {
return false
}
...
}
```
The session cache key is `"auth:session:" + token` with `token` taken from the
`Authorization` header, so its length and byte content are chosen by the client.
**Failure scenario.** A deployment uses memcache and issues JWT session tokens,
which routinely exceed 238 bytes. `MemcacheProvider.Get` receives
`ErrMalformedKey` and — per finding 17 — reports it as a plain cache miss
(`provider_memcache.go:80-82`). The loader runs, the database validates the
session, and then `MemcacheProvider.Set` returns `ErrMalformedKey`, which
finding 1 converts into an authentication failure. **Every user with a long
token is permanently unable to authenticate**, and the logs show nothing but a
generic "failed to cache value". A client can also trigger this deliberately
with a token containing a space to probe the backend.
**Recommendation.** Hash keys inside the provider so key length and charset are
bounded regardless of caller input — e.g. `sha256` hex of the key when it
exceeds 200 bytes or contains illegal bytes, with a fixed prefix. And, as in
finding 7, `pkg/security` should hash the token before it ever becomes a key.
---
### 10. Medium — TOCTOU when deleting an expired item
`provider_memory.go:66-74`:
```go
if item.isExpired() {
m.mu.RUnlock()
// Upgrade to write lock to delete expired item
m.mu.Lock()
delete(m.items, key)
m.mu.Unlock()
m.misses.Add(1)
return nil, false
}
```
Go's `RWMutex` has no lock upgrade, so the read lock is genuinely released and
the write lock separately acquired. The `delete` is then **unconditional** — it
does not re-check that the entry still exists or is still the expired one.
**Failure scenario.** Goroutine A reads an expired entry for key `K` and drops
the read lock. Goroutine B takes the write lock and `Set`s a fresh value for
`K`. Goroutine A now takes the write lock and deletes B's fresh entry. B has no
idea; the value it believes it cached is gone, and the next reader recomputes
it. In the auth path this means a just-validated session is dropped immediately,
forcing another stored-procedure call — and under load this can repeat, since the
interleaving recurs whenever expiry and refresh coincide, which is exactly when
traffic for that key is highest.
**Recommendation.** Re-check under the write lock and delete only the same
entry:
```go
m.mu.Lock()
if cur, ok := m.items[key]; ok && cur == item {
m.removeLocked(key) // see finding 4
}
m.mu.Unlock()
```
---
### 11. Medium — `Close()` makes the provider panic on next use
`provider_memory.go:273-280`:
```go
func (m *MemoryProvider) Close() error {
m.mu.Lock()
defer m.mu.Unlock()
m.items = nil
return nil
}
```
Reads of a nil map are fine, but `Set`/`SetWithTags` do
`m.items[key] = &memoryItem{...}` (`:111`, `:154`), which on a nil map panics
with `assignment to entry in nil map`.
**Failure scenario.** A graceful-shutdown handler calls `cache.Close()` while
requests are still draining — the normal ordering, since `pkg/config` defaults
`servers.drain_timeout` to 25s. The next in-flight request reaches
`setQueryTotalCache` and the process panics **while holding `m.mu`**. There is
no `recover()` anywhere in `pkg/cache`, so this unwinds into whatever the caller
has; if the HTTP server's panic handler recovers it, the mutex is never
unlocked and **every subsequent cache operation blocks forever** — the process
is alive, serving, and permanently wedged on the first cache access. That is a
worse outcome than crashing.
**Recommendation.** Mark the provider closed instead of destroying the map, and
return an error from every method afterwards:
```go
type MemoryProvider struct {
...
closed bool
}
func (m *MemoryProvider) Close() error {
m.mu.Lock()
defer m.mu.Unlock()
m.closed = true
m.items = make(map[string]*memoryItem)
m.tagToKeys = make(map[string]map[string]struct{})
return nil
}
```
with `if m.closed { return ErrClosed }` at the top of each mutating method.
More generally, the package should use `defer logger.CatchPanicCallback(...)` at
its exported boundaries so a panic inside a lock is reported rather than
silently converted into a deadlock — but note `audit/pkg/logger.audit.md`
finding 4 on `CatchPanic` swallowing unconditionally; here the panic should be
logged and **re-raised** or converted to an error, not absorbed.
---
### 12. Medium — the memcache tag index is a lost-update machine
`provider_memcache.go:136-170` maintains, for each tag, a JSON array of keys:
```go
for _, tag := range tags {
tagKey := fmt.Sprintf("cache:tag:%s", tag)
// Get existing keys for this tag
var keys []string
if item, err := m.client.Get(tagKey); err == nil {
_ = json.Unmarshal(item.Value, &keys)
}
// Add current key if not already present
found := false
for _, k := range keys { ... }
if !found {
keys = append(keys, key)
}
keysData, err := json.Marshal(keys)
if err != nil {
continue
}
tagItem := &memcache.Item{
Key: tagKey,
Value: keysData,
Expiration: expiration + 3600, // Give tag lists longer TTL
}
_ = m.client.Set(tagItem)
}
```
Four distinct defects in fifteen lines:
1. **Non-atomic read-modify-write.** `Get` then `Set` with no CAS, across a
network, for a value that every concurrent writer of the same tag touches.
2. **Unbounded growth.** The list only ever grows within its TTL, and every
writer re-reads and re-writes the whole array. `Delete` (`:190-200`) rewrites
it too.
3. **1 MB item limit.** Once the array exceeds memcached's default item size the
`Set` fails — and the error is discarded by `_ =`.
4. **`expiration + 3600` can cross memcached's 30-day boundary.** The protocol
treats an expiry value above 2 592 000 as an **absolute Unix timestamp**. A
caller passing a 30-day TTL yields `2592000 + 3600 = 2595600`, which
memcached reads as 1970-01-31 — already past, so the tag list is dead on
arrival. `int32(ttl.Seconds())` at `:108` also truncates for very large TTLs
and, for a negative TTL, expires the item immediately.
**Failure scenario.** Two requests cache different query totals for the same
table concurrently. Both read `cache:tag:table:orders` and see `[k1]`; one
writes `[k1,k2]`, the other writes `[k1,k3]`. The second write wins and `k2` is
**no longer in the tag index**. A later `POST` to that table calls
`invalidateCacheForTags` → `DeleteByTag("table:orders")`, which deletes `k1` and
`k3` but not `k2`. `k2` continues to serve a **stale row count from before the
write** for its full TTL. Under concurrency this is not an edge case — it is the
normal outcome, and lost updates accumulate. When the list crosses 1 MB the
failure becomes total: no further keys are indexed and nothing is logged.
Additionally, the prefixes `cache:tag:` and `cache:tags:` (`:128`, `:138`,
`:179`, `:185`, `:219`, `:239`) share the key namespace with ordinary cache
keys. No consumer currently writes keys under that prefix — the restheadspec
keys are `query_total:<sha256>` and the security keys `auth:session:*` — but
nothing enforces it, and a consumer that ever allows an attacker-derived key
beginning `cache:tag:` could forge or destroy the invalidation index (mass
invalidation → database load, or suppressed invalidation → stale authorization
data served).
**Recommendation.** Use `CompareAndSwap` (gomemcache exposes it) in a bounded
retry loop, cap the list length, honour the 30-day rule, and stop discarding
errors:
```go
func memcacheExpiry(ttl time.Duration) int32 {
secs := int64(ttl.Seconds())
if secs < 0 { secs = 0 }
if secs > 2592000 { // >30d must be an absolute timestamp
return int32(time.Now().Add(ttl).Unix())
}
return int32(secs)
}
```
Given how weak tag support is on memcache, the honest alternative is to return
`ErrUnsupported` from `SetWithTags`/`DeleteByTag` and force callers to pick a
provider that can do it, rather than offering invalidation that silently misses
keys. Namespace the index keys under a prefix that ordinary keys cannot reach
(e.g. by prefixing all user keys with `k:`).
---
### 13. Medium — `Clear()` flushes the entire shared server
```go
// provider_memcache.go:257-259
func (m *MemcacheProvider) Clear(ctx context.Context) error {
return m.client.FlushAll()
}
// provider_redis.go:228-230
func (r *RedisProvider) Clear(ctx context.Context) error {
return r.client.FlushDB(ctx).Err()
}
```
Neither is scoped to this application's keys. `FlushAll` wipes every key on
every configured memcached server; `FlushDB` wipes the whole logical Redis DB —
and `RedisConfig.DB` defaults to `0`, which is also where `pkg/config` puts the
event broker (`event_broker.redis.db: 0`, `pkg/config/manager.go:265`) and its
`resolvespec:events` stream.
**Failure scenario.** An operator or an admin endpoint calls `cache.Clear()` to
drop stale entries. On Redis with the default config this also deletes the event
broker's stream and consumer-group state, so queued events are lost and
consumers fail; on memcached it evicts every other tenant sharing that server.
There is no confirmation, no scoping, and the method is one call away from any
consumer holding the `*Cache`.
**Recommendation.** Implement `Clear` as a scoped delete over the application's
own key prefix (`SCAN`+`DEL` for Redis), require an explicitly opted-in
"destructive flush" flag to use `FlushDB`/`FlushAll`, and document that the
cache must not share a Redis DB with the event broker. Consider adding a
mandatory `KeyPrefix` to `Options` so scoping is always possible.
---
### 14. Medium — `MemcacheProvider.Close()` is a no-op based on a false comment
`provider_memcache.go:267-271`:
```go
func (m *MemcacheProvider) Close() error {
// Memcache client doesn't have a close method
return nil
}
```
This is factually wrong for the pinned dependency: gomemcache
`v0.0.0-20260422231931-4d751bb6e37c` exposes `func (c *Client) Close() error` at
`memcache.go:836`. Its documented behaviour is to close all currently-open idle
connections (the client stays usable afterwards), which is exactly what a
provider `Close()` should be releasing.
**Failure scenario.** Every provider replacement (finding 3) and every
`cache.Close()` leaves the memcache client's idle connections — up to
`MaxIdleConns` per server — established. In a process that reconfigures the
cache, or in tests that construct providers repeatedly, file descriptors
accumulate until `accept`/`dial` starts failing with `too many open files`,
which manifests as unrelated failures elsewhere in the process.
**Recommendation.** `return m.client.Close()`. Also add `Close` handling to the
swap paths in `cache.go` per finding 3.
---
### 15. Medium — no TLS option for Redis or memcache
`RedisConfig` (`provider_redis.go:17-36`) has `Host`, `Port`, `Password`, `DB`,
`PoolSize`, `Options` — and nothing else. `redis.Options` supports `TLSConfig`,
but the constructor never sets it:
```go
client := redis.NewClient(&redis.Options{
Addr: fmt.Sprintf("%s:%d", config.Host, config.Port),
Password: config.Password,
DB: config.DB,
PoolSize: config.PoolSize,
})
```
`MemcacheConfig` (`:18-31`) likewise has no transport security, and the
memcached protocol has no in-band auth here at all.
**Failure scenario.** The cached values are `UserContext` objects — identity and
authorization data — and the `AUTH` password is sent in cleartext on the first
command of every new connection. Anything able to observe the path between the
service and Redis (a shared VPC, a misconfigured security group, a compromised
sidecar, a managed Redis reached over the public internet) can read session
contents and steal the Redis password, then write forged `auth:session:*` entries
directly. Writing a crafted `UserContext` into the cache is a **complete
authentication bypass**: `GetOrSet` returns it on a hit and never consults the
database.
This is the same theme as `pkg/config`'s `sslmode: disable` default
(`pkg/config/manager.go:242`) — see `audit/pkg/config.audit.md` finding 3.
**Recommendation.** Add TLS configuration to both configs and plumb it through:
```go
type RedisConfig struct {
...
TLS bool
TLSSkipVerify bool // must default false
TLSCACertFile string
}
```
and set `redis.Options.TLSConfig` accordingly. Since the cache holds
authentication material, treat encrypted transport as the default and require an
explicit opt-out. For memcache, prefer Redis for this workload or terminate TLS
with a local proxy (stunnel/envoy) and document it.
---
### 16. Medium — `Remember` returns a different type on hit vs miss
`cache_manager.go:145-167`:
```go
func (c *Cache) Remember(ctx context.Context, key string, ttl time.Duration, loader func() (interface{}, error)) (interface{}, error) {
data, err := c.GetBytes(ctx, key)
if err == nil {
var result interface{}
if err := json.Unmarshal(data, &result); err == nil {
return result, nil // <-- map[string]interface{} / float64 / ...
}
}
value, err := loader()
...
return value, nil // <-- whatever the loader returned
}
```
On a hit the value comes back as generic JSON (`map[string]interface{}`,
`[]interface{}`, `float64`, `string`); on a miss it is the loader's concrete Go
type.
**Failure scenario.** A caller writes the natural thing:
```go
v, err := c.Remember(ctx, key, ttl, func() (any, error) { return loadUser(id) })
u := v.(*User) // panics on every cache hit
```
This passes every test run against a cold cache and panics in production as soon
as the cache warms — the worst possible failure timing. `Remember` has no
callers in `pkg/` today, so this is latent, but it is a trap laid for the next
consumer and there is no `recover()` in the package to contain it.
**Recommendation.** Either delete `Remember` in favour of `GetOrSet` (which
takes a typed `dest`), or give it the same contract:
```go
func (c *Cache) Remember(ctx context.Context, key string, dest any, ttl time.Duration, loader func() (any, error)) error
```
If the generic-return shape must stay, document it loudly and name it
`RememberAny`.
---
### 17. Medium — `Get`/`Exists` swallow every backend error as a cache miss
```go
// provider_redis.go:86-95
val, err := r.client.Get(ctx, key).Bytes()
if err == redis.Nil {
return nil, false
}
if err != nil {
return nil, false // network error, WRONGTYPE, auth failure, timeout...
}
// provider_memcache.go:75-84 — same shape
// provider_memcache.go:262-265
func (m *MemcacheProvider) Exists(ctx context.Context, key string) bool {
_, err := m.client.Get(key)
return err == nil
}
```
The `Provider.Get` signature `([]byte, bool)` cannot express "I don't know", so
every failure is indistinguishable from an absent key. Nothing is logged —
`pkg/cache` imports no logger.
**Failure scenario.** Redis develops packet loss, or memcached is restarted, or
`maxmemory` is hit. Every lookup reports a miss, so every request falls through
to the database: the count query at `restheadspec/handler.go:857` and the session
stored procedure at `providers.go:413`. Load multiplies by the cache hit ratio
in an instant — typically 10–100× — and the database becomes the next thing to
fall over. Meanwhile the metrics say "cache miss", the logs say nothing, and the
`Stats` endpoint reports a plausible-looking miss count, so the actual cause is
invisible during the incident. Combined with finding 1, the subsequent `Set`
failure also rejects every request, so the symptom presented to operators is
"authentication is broken" with no mention of Redis.
**Recommendation.** Widen the interface to `Get(ctx, key) ([]byte, bool, error)`
(or keep the two-value form and add `GetE`), and at minimum log at warn level
with a rate limit inside each provider:
```go
if err != nil && err != redis.Nil {
logger.Warn("cache: redis GET failed for prefix %s: %v", ctx, keyPrefix(key), err)
return nil, false
}
```
Note `keyPrefix` rather than `key`, per finding 7. Feed a
`cache_errors_total{provider}` counter into `pkg/metrics` so a degraded backend
is alertable — the `Provider` interface there already has `RecordCacheHit`/
`RecordCacheMiss` but no error counter.
---
### 18. Low — `evictOne` is an O(n) scan under the write lock
`provider_memory.go:306-326`:
```go
func (m *MemoryProvider) evictOne() {
var oldestKey string
var oldestTime time.Time
for key, item := range m.items {
if item.isExpired() {
delete(m.items, key)
return
}
if oldestKey == "" || item.LastAccess.Before(oldestTime) {
oldestKey = key
oldestTime = item.LastAccess
}
}
if oldestKey != "" {
delete(m.items, oldestKey)
}
}
```
Called from `Set` (`:107`) and `SetWithTags` (`:137`) whenever the cache is at
capacity — i.e. on **every insertion** in steady state — while the write lock is
held. With the default `MaxSize: 10000` that is a 10 000-entry map walk plus
10 000 `time.Time` comparisons per write, blocking all readers (which, per
finding 5, includes every cache hit).
`Stats` (`:283-304`) has the same O(n) shape under `RLock`, and its comment
"Clean expired items first" is wrong — it holds a **read** lock and only counts.
**Recommendation.** Use a proper LRU (an intrusive doubly-linked list beside the
map, as `hashicorp/golang-lru` or `container/list` gives you) for O(1) eviction,
or sample k random entries and evict the oldest of those (Redis's approach) if
approximate LRU is acceptable. Maintain a counter for `Stats` instead of
walking. Fix the misleading comment.
---
### 19. Low — `CleanExpired` is dead code; there is no janitor
`provider_memory.go:329` `CleanExpired` has no callers anywhere in the
repository (verified by grep — the only hits are its own declaration and
comment). No goroutine sweeps expirations.
**Failure scenario.** Expired entries are reclaimed only when someone looks them
up (`Get`, `:66`) or when `evictOne` happens to walk past one. A workload whose
keys are written once and never re-read — which is exactly the
`query_total:<sha256>` pattern, since a distinct filter combination is usually
requested once — retains every expired entry until `MaxSize` forces eviction.
The cache therefore sits permanently at its maximum footprint holding mostly
dead data, and the LRU scan in finding 18 walks those dead entries on every
insertion. With `MaxSize <= 0` (finding 22) nothing reclaims them at all.
**Recommendation.** Start a janitor goroutine from `NewMemoryProvider` with a
configurable interval, and stop it in `Close()`:
```go
func NewMemoryProvider(opts *Options) *MemoryProvider {
m := &MemoryProvider{...; done: make(chan struct{})}
go m.janitor(opts.CleanupInterval) // default e.g. 1 minute
return m
}
```
Guard the goroutine with `defer logger.CatchPanicCallback("cache.janitor", ...)`
so a panic in the sweep is reported rather than killing the process silently.
---
### 20. Low — `RedisProvider.Stats` returns the raw `INFO` output
`provider_redis.go:247-268`:
```go
info, err := r.client.Info(ctx, "stats", "keyspace").Result()
...
stats := &CacheStats{
Keys: dbSize,
ProviderType: "redis",
ProviderStats: map[string]any{
"info": info,
},
}
```
`CacheStats.ProviderStats` is `json:"provider_stats,omitempty"` — i.e. designed
to be serialized. The `stats` and `keyspace` sections include the Redis version,
uptime, connected-client counts, keyspace hit/miss totals, eviction and
expiration counters, and per-database key counts.
**Failure scenario.** `cache.GetStats` (`cache.go:65`) has no callers today, but
it is an obvious thing to wire to a `/health` or `/admin/stats` endpoint. Doing
so exposes infrastructure fingerprinting to any client that reaches it, and the
keyspace counters leak activity volume. It is an information-disclosure
primitive waiting for a route.
**Recommendation.** Parse `INFO` into a small allowlisted set of numeric fields
(`keyspace_hits`, `keyspace_misses`, `evicted_keys`, `expired_keys`) and
populate `CacheStats.Hits`/`Misses` from them rather than passing the blob
through. If the raw text is wanted for debugging, gate it behind an explicit
debug flag and never include it in a response body.
---
### 21. Low — the memcache provider ignores `ctx` entirely
Every method on `MemcacheProvider` accepts `ctx context.Context` and none uses
it (`:75`, `:87`, `:103`, `:177`, `:218`, `:252`, `:257`, `:262`, `:275`). The
only timeout is the client-wide `config.Timeout` (default 1s, `:49-51`).
**Failure scenario.** A client disconnects or the request deadline expires; the
handler's context is cancelled, but the cache call continues to completion. In
`SetWithTags` that is one `Set` plus, per tag, a `Get` and a `Set` — so a
two-tag write is five sequential round trips, each able to consume the full 1s
client timeout, all after the caller has given up. Under load-shedding
conditions the server keeps doing work for requests nobody is waiting for, which
is precisely when it can least afford to.
**Recommendation.** gomemcache's API is context-free, so either wrap each call
with a `select` on `ctx.Done()` and a goroutine, or switch to a
context-aware client. At minimum, check `ctx.Err()` at the top of each method
and return early, and document that `config.Timeout` is the real bound.
---
### 22. Low — `MaxSize <= 0` silently disables eviction
`provider_memory.go:105` and `:135` both guard with `if m.options.MaxSize > 0 && ...`.
A caller who constructs `&Options{DefaultTTL: time.Minute}` — as
`NewRedisProvider` (`:59`) and `NewMemcacheProvider` (`:54`) do for their own
defaults, and as any hand-written config easily does — gets an **unbounded**
in-memory cache. Combined with findings 4 and 19, memory then grows with
attacker-controlled query variation until the process is OOM-killed.
**Recommendation.** Treat `MaxSize <= 0` as "use the default" (10 000) rather
than "unlimited", and require an explicit sentinel such as `MaxSize: -1` to
opt into unbounded. Validate `Options` in `NewMemoryProvider` and log the
effective values once at startup.
---
### 23. Low — constructors mutate the caller's config struct
`NewMemcacheProvider` writes `config.Servers`, `config.MaxIdleConns`,
`config.Timeout` and `config.Options` (`:41-57`); `NewRedisProvider` writes
`config.Host`, `config.Port`, `config.PoolSize`, `config.Options` (`:48-62`).
Both also assign into the `config == nil` replacement, which is at least local.
**Failure scenario.** A caller holds one config struct and constructs two
providers from it (a common test pattern, or a primary/replica setup). The second
construction sees the first one's defaults already applied, so "zero means
default" no longer holds and an intentional later change is silently ignored.
Callers reasonably assume a constructor does not write to their arguments.
**Recommendation.** Copy first: `cfg := *config` (plus a copy of `Options` if it
is non-nil, since it is a pointer), then apply defaults to the local copy.
---
### 24. Low — no defensive copy of `[]byte` in the memory provider
`Set` stores the caller's slice directly (`provider_memory.go:112`) and `Get`
returns the stored slice directly (`:77`, `:87`). Both share backing memory with
the caller.
**Failure scenario.** Two requests `GetBytes` the same key and receive the same
underlying array. If either mutates it in place — or if a caller reuses a buffer
it passed to `SetBytes` — the cached value changes for everyone, with no lock
held and no copy. For the session cache that means one request's scratch buffer
can rewrite another user's cached `UserContext`. Today every consumer goes
through `json.Marshal`/`json.Unmarshal` in `cache_manager.go`, which allocates
fresh slices, so this is latent rather than live; the `SetBytes`/`GetBytes`
API (`:56`, `:37`) exposes it directly to any future caller.
**Recommendation.** Copy on both sides in `MemoryProvider` — the Redis and
memcache providers get copies for free because the data crosses a socket, so
this also removes a behavioural difference between providers:
```go
buf := make([]byte, len(value))
copy(buf, value)
```
Document the ownership rule on the `Provider` interface either way.
---
### 25. Low — `example_usage.go` is a library file full of `log.Fatal`
`pkg/cache/example_usage.go` (266 lines) is compiled into the package and calls
`log.Fatal` at seventeen sites (`:18`, `:36`, `:44`, `:57`, `:64`, `:83`, `:96`,
`:103`, `:112`, `:127`, `:146`, `:154`, `:168`, `:187`, `:199`, `:224`, `:245`).
`log.Fatal` calls `os.Exit(1)`.
**Failure scenario.** These are exported functions (`ExampleInMemoryCache`,
`ExampleRedisCache`, `ExampleMemcacheCache`, …) with no `_test.go` suffix, so
they are part of the package's public API. Anything that calls one — a
misremembered name, a code-completion accident, a copied snippet — can terminate
the host process on a cache error, bypassing every panic handler and graceful
shutdown path. They also drag `log` into the package's dependency set while the
package deliberately imports no logger.
**Recommendation.** Move the file to `example_usage_test.go` (Go's testable-example
convention, which also makes them compile-checked and runnable), or to a
`_examples/` directory outside the package. Replace `log.Fatal` with returned
errors.
---
## What looks right
- `Provider` (`provider.go:9-44`) is a clean, minimal interface; the three
backends are genuinely swappable and the package has no import cycle problems
(it imports nothing from `ResolveSpec` at all).
- `hits`/`misses` use `atomic.Int64` (`provider_memory.go:35-36`) and are read
with `Load()` in `Stats`, so the counters themselves are race-free.
- `MemoryProvider.Delete` (`:173-191`) and `SetWithTags` (`:141-151`) do maintain
`tagToKeys` correctly, including deleting the tag entry when its key set
empties — the bug in finding 4 is the *other* paths not doing the same.
- `DeleteByTag` (`:194-228`) correctly handles multi-tag items: it strips only
the invalidated tag and keeps the item alive if other tags remain.
- `RedisProvider.DeleteByPattern` (`:196-225`) batches `DEL` in groups of 100
rather than buffering an unbounded pipeline, and checks `iter.Err()`.
- `RedisProvider.Close` (`:242-244`) correctly delegates to `client.Close()`.
- Both network providers verify connectivity at construction
(`provider_redis.go:75`, `provider_memcache.go:64`) and return an error rather
than deferring the failure to the first request — though the Redis `Ping` uses
a 5s blocking timeout on a `context.Background()`, which delays startup.
- Cache keys for query totals are SHA-256 hashes of the query shape
(`restheadspec/cache_helpers.go:88-92`), not raw user input — which is why the
key-injection risk in findings 9 and 12 is confined to the `pkg/security`
session path.
- `pkg/security/keystore_database.go:287` already keys on a hash
(`keystoreCacheKey(hash)`), which is the pattern `providers.go:398` should
follow.
## Suggested follow-up
Ordered by value:
1. **Make cache failures non-fatal in the auth path** (findings 1, 2, 9). This
is the single highest-value change: a cache problem should never be able to
reject a valid credential, and a revocation that cannot be honoured must be
loud.
2. **Key the session cache on `sha256(token)`** (findings 7, 9, 12) in
`pkg/security/providers.go:318`, `:398`, `:466`. Removes attacker control of
cache keys, bounds their length, and keeps the credential out of error
strings.
3. **Replace `DeleteByPattern`-based revocation with `DeleteByTag`**
(findings 2, 6), and make `ClearUserCache` actually scoped to the user.
4. **Fix the `defaultCache` race** (finding 3) with `atomic.Pointer` +
`sync.Once`, and close the displaced provider on swap (finding 14).
5. **Add `go test -race ./pkg/cache/...` to CI.** Findings 3 and 10 are
detectable in minutes with a small concurrent test; see
`_CROSS-CUTTING.audit.md`. `pkg/cache` currently has 69 lines of tests — two
functions, `TestSetDefaultCache` (`cache_test.go:9`) and
`TestGetDefaultCacheInitialization` (`:50`) — and is not in the package list
that `Makefile`/`.github/workflows/tests.yml` run.
6. **Centralize tag-index maintenance in `MemoryProvider`** (finding 4) and reset
it in `Clear`; add an eviction path that cannot leak.
7. **Make `MemoryProvider.Get` read-only** (finding 5) by moving `LastAccess`
and `HitCount` to atomics.
8. **Add single-flight to `GetOrSet`** (finding 8).
9. **Add TLS to `RedisConfig`/`MemcacheConfig`** (finding 15) and default it on.
10. **Give the package a logger and error metrics** (finding 17). Right now
`pkg/cache` cannot tell anyone that anything went wrong: 1 538 lines, zero
log statements, zero `recover()`, and eight `_ =` error discards outside the
examples file — seven of them in `provider_memcache.go`, precisely where the
tag index breaks (finding 12).
11. **Scope `Clear`** (finding 13) so it cannot flush the event broker's Redis DB.
## Cross-references
- `audit/pkg/logger.audit.md` — finding 2 (messages forwarded to Sentry
unscrubbed) is what makes finding 7 here dangerous; finding 4 (`CatchPanic`
swallows) is why findings 11/16/19 should not simply add `CatchPanic`.
- `audit/pkg/config.audit.md` — finding 3 (insecure transport defaults) is the
same theme as finding 15; `cache.redis.*` defaults live at
`pkg/config/manager.go:195-202` and expose no TLS field.
- `audit/pkg/security.audit.md` — findings 1, 2, 7, 8, 9 all land in
`pkg/security/providers.go`; the unbounded `go a.updateSessionActivity(...)`
per request (`:447`) is raised there.
- `audit/pkg/restheadspec.audit.md` — the query-total cache path
(`handler.go:829-860`) and tag-based invalidation (`:1477`, `:1711`, `:1785`,
`:1859`, `:1919`, `:2024`) are the other consumer; error returns from
`invalidateCacheForTags` are the invalidation-failure signal.
- `audit/pkg/metrics.audit.md` — `Provider` has `RecordCacheHit`/
`RecordCacheMiss`/`UpdateCacheSize` but no cache-error counter, and
`pkg/cache` calls none of them.
- `audit/pkg/_CROSS-CUTTING.audit.md` — no `-race` in CI; only
`pkg/resolvespec` and `pkg/restheadspec` are tested.