mirror of
https://github.com/bitechdev/ResolveSpec.git
synced 2026-10-01 04:21:58 +00:00
Tests / Unit Tests (push) Failing after 24s
Tests / Integration Tests (push) Failing after 26s
Build , Vet Test, and Lint / Build (push) Successful in 1m7s
Build , Vet Test, and Lint / Run Vet Tests (1.23.x) (push) Successful in 1m31s
Build , Vet Test, and Lint / Lint Code (push) Successful in 1m33s
Build , Vet Test, and Lint / Run Vet Tests (1.24.x) (push) Successful in 1m34s
1360 lines
56 KiB
Markdown
1360 lines
56 KiB
Markdown
# Audit: `pkg/cache`
|
||
|
||
| | |
|
||
|---|---|
|
||
| **Package** | `github.com/bitechdev/ResolveSpec/pkg/cache` |
|
||
| **Files** | `cache.go` (76), `cache_manager.go` (167), `provider.go` (65), `provider_memory.go` (342), `provider_memcache.go` (284), `provider_redis.go` (269), `example_usage.go` (266), `cache_test.go` (69) |
|
||
| **Audit date** | 2026-09-29 |
|
||
| **Axes** | thread locking/waiting, slowness, security, panic handling & logging |
|
||
| **Threat model** | hostile internet client; request bodies, headers, query params, schema/table/column names all attacker-controlled |
|
||
| **Depth** | deep (hot package) |
|
||
|
||
## Summary
|
||
|
||
`pkg/cache` is a thin `Provider` abstraction over three backends. Two things
|
||
make it much more security-relevant than a cache normally is:
|
||
|
||
1. **It stores authentication state.** `pkg/security/providers.go:398-402` caches
|
||
`UserContext` under the key `fmt.Sprintf("auth:session:%s", token)` — the raw
|
||
bearer token from the `Authorization` header. So **cache keys are directly
|
||
attacker-controlled**, the cached value is an authorization decision, and
|
||
`DeleteByPattern` is the only session-revocation mechanism.
|
||
2. **It is never explicitly initialized.** Nothing in `pkg/` calls
|
||
`Initialize`/`UseMemory`/`UseRedis`/`UseMemcache`; every consumer reaches the
|
||
cache through `GetDefaultCache()` (`cache.go:48`), which lazily constructs a
|
||
`MemoryProvider` **on the request path** without synchronization.
|
||
|
||
The most serious findings are: a cache-write failure being converted into an
|
||
authentication failure (`GetOrSet`, `cache_manager.go:126`), session revocation
|
||
that silently cannot work on memcache, an unsynchronized lazy singleton read by
|
||
concurrent HTTP handlers, an unbounded `tagToKeys` index that `MaxSize` does not
|
||
cover, and `MemoryProvider.Get` taking the **write** lock on every cache hit.
|
||
|
||
There is **no `recover()` anywhere in this package** and no logging at all —
|
||
grep for `recover()` and `logger.` across all 1 538 lines returns zero hits.
|
||
Every error is either returned to the caller or discarded with `_ =`.
|
||
|
||
## Findings
|
||
|
||
| # | Severity | Axis | Finding |
|
||
|---|---|---|---|
|
||
| 1 | **Critical** | security | Cache-write failure is returned as an error from `GetOrSet`, so a cache outage becomes a total authentication outage |
|
||
| 2 | **Critical** | security | Session revocation (`DeleteByPattern`) is unimplemented on memcache and returns an error — revoked sessions keep authenticating |
|
||
| 3 | **High** | locking | `defaultCache` is an unsynchronized lazy singleton read/written from concurrent request handlers |
|
||
| 4 | **High** | security / slowness | `tagToKeys` is unbounded and leaked by six paths; `MaxSize` bounds only `items` |
|
||
| 5 | **High** | locking / slowness | `MemoryProvider.Get` acquires the **write** lock on every hit |
|
||
| 6 | **High** | security | `DeleteByPattern` pattern syntax differs per provider (Go regexp vs Redis glob vs error); memory matches unanchored |
|
||
| 7 | **High** | security | Raw session token embedded in the `"key not found: %s"` error string |
|
||
| 8 | **High** | security | No single-flight: concurrent misses on one key run N loaders (auth stampede against the session stored procedure) |
|
||
| 9 | **Medium** | security | Attacker-controlled keys break memcache's 250-byte/no-whitespace key rule; error is swallowed as a miss, then fails the write |
|
||
| 10 | **Medium** | correctness | TOCTOU in `MemoryProvider.Get`: expired-item path deletes unconditionally after dropping the read lock |
|
||
| 11 | **Medium** | correctness | `MemoryProvider.Close()` sets `items = nil`; a subsequent `Set` panics and nothing recovers |
|
||
| 12 | **Medium** | security | Memcache tag index is non-atomic read-modify-write and shares the key namespace; lost updates silently drop keys from invalidation |
|
||
| 13 | **Medium** | security | `Clear()` maps to `FlushAll()` / `FlushDB()` — wipes the whole shared server/DB |
|
||
| 14 | **Medium** | correctness | `MemcacheProvider.Close()` is a no-op justified by a false comment; gomemcache **does** have `Close()` |
|
||
| 15 | **Medium** | security | Redis and memcache have no TLS option at all; `AUTH` password and cached `UserContext` cross the wire in cleartext |
|
||
| 16 | **Medium** | correctness | `Cache.Remember` returns a different Go type on hit vs miss |
|
||
| 17 | **Medium** | slowness | `Get`/`Exists` swallow **all** backend errors as a cache miss — a degraded backend is invisible and stampedes the DB |
|
||
| 18 | **Low** | slowness | `evictOne` is an O(n) scan under the write lock, run per insertion at capacity |
|
||
| 19 | **Low** | correctness | `CleanExpired` is dead code; there is no janitor, so expired items are only reclaimed on access |
|
||
| 20 | **Low** | security | `RedisProvider.Stats` returns the raw `INFO` output in `ProviderStats["info"]` |
|
||
| 21 | **Low** | correctness | `ctx` is accepted and completely ignored by the memcache provider |
|
||
| 22 | **Low** | correctness | `MaxSize <= 0` disables eviction entirely — an unbounded in-memory cache |
|
||
| 23 | **Low** | correctness | Provider constructors mutate the caller's config struct |
|
||
| 24 | **Low** | correctness | No defensive copy of `[]byte` on `Set`/`Get` in the memory provider |
|
||
| 25 | **Low** | hygiene | `example_usage.go` ships `log.Fatal` calls in a library package |
|
||
|
||
---
|
||
|
||
### 1. Critical — a cache-write failure is an authentication failure
|
||
|
||
`pkg/cache/cache_manager.go:112-141`:
|
||
|
||
```go
|
||
func (c *Cache) GetOrSet(ctx context.Context, key string, dest interface{}, ttl time.Duration, loader func() (interface{}, error)) error {
|
||
err := c.Get(ctx, key, dest)
|
||
if err == nil {
|
||
return nil
|
||
}
|
||
|
||
value, err := loader()
|
||
if err != nil {
|
||
return fmt.Errorf("loader failed: %w", err)
|
||
}
|
||
|
||
// Store in cache
|
||
if err := c.Set(ctx, key, value, ttl); err != nil {
|
||
return fmt.Errorf("failed to cache value: %w", err) // <-- line 127
|
||
}
|
||
...
|
||
}
|
||
```
|
||
|
||
The loader has already succeeded — the authoritative value is in hand — but a
|
||
failure to *cache* it aborts the whole call. The only consumer of `GetOrSet` is
|
||
authentication, `pkg/security/providers.go:398-444`:
|
||
|
||
```go
|
||
cacheKey := fmt.Sprintf("auth:session:%s", token)
|
||
|
||
var userCtx UserContext
|
||
err := a.cache.GetOrSet(r.Context(), cacheKey, &userCtx, a.cacheTTL, func() (any, error) {
|
||
// ... queries the session stored procedure, returns &user on success
|
||
})
|
||
|
||
if err != nil {
|
||
lastErr = err
|
||
continue // Try next token
|
||
}
|
||
```
|
||
|
||
Any error — including "failed to cache value" — is treated as *this token is not
|
||
valid*, and after the token loop the request is rejected.
|
||
|
||
**Failure scenario.** Redis is configured and becomes unreachable (restart,
|
||
failover, network partition, `maxmemory` reached with `noeviction`).
|
||
`RedisProvider.Get` swallows the error and reports a miss
|
||
(`provider_redis.go:91-93`), the loader runs and the database confirms the
|
||
session is valid, then `RedisProvider.Set` returns the connection error and
|
||
`GetOrSet` returns it. **Every request from every user is now rejected with an
|
||
authentication error**, even though both the database and the sessions are
|
||
healthy. A cache is supposed to be a latency optimization; here it is a hard
|
||
dependency of the auth path, and its failure mode is total outage. The same
|
||
applies to `maxmemory` pressure, which an attacker can induce (see finding 4).
|
||
|
||
**Recommendation.** A cache-write failure must be non-fatal. Log it and return
|
||
the loaded value:
|
||
|
||
```go
|
||
if err := c.Set(ctx, key, value, ttl); err != nil {
|
||
logger.Warn("cache: failed to store key (continuing uncached): %v", ctx, err)
|
||
}
|
||
```
|
||
|
||
Separately, `pkg/security` should not conflate "cache layer failed" with
|
||
"credential rejected"; the loader's own error is the only one that should fail
|
||
authentication. Consider having `GetOrSet` return the loaded value plus a
|
||
non-fatal cache error, or wrap cache errors in a sentinel the caller can test
|
||
with `errors.Is`.
|
||
|
||
---
|
||
|
||
### 2. Critical — session revocation silently cannot work on memcache
|
||
|
||
`pkg/security/providers.go:463-479` is the only session-revocation path:
|
||
|
||
```go
|
||
func (a *DatabaseAuthenticator) ClearCache(token string) error {
|
||
ctx := context.Background()
|
||
if token != "" {
|
||
cacheKey := fmt.Sprintf("auth:session:%s", token)
|
||
return a.cache.Delete(ctx, cacheKey)
|
||
}
|
||
// Clear all auth cache entries
|
||
return a.cache.DeleteByPattern(ctx, "auth:session:*")
|
||
}
|
||
|
||
func (a *DatabaseAuthenticator) ClearUserCache(userID int) error {
|
||
ctx := context.Background()
|
||
pattern := "auth:session:*"
|
||
return a.cache.DeleteByPattern(ctx, pattern)
|
||
}
|
||
```
|
||
|
||
`pkg/cache/provider_memcache.go:249-254`:
|
||
|
||
```go
|
||
// DeleteByPattern removes all keys matching the pattern.
|
||
// Note: Memcache does not support pattern-based deletion natively.
|
||
// This is a no-op for memcache and returns an error.
|
||
func (m *MemcacheProvider) DeleteByPattern(ctx context.Context, pattern string) error {
|
||
return fmt.Errorf("pattern-based deletion is not supported by Memcache")
|
||
}
|
||
```
|
||
|
||
**Failure scenario.** A deployment uses memcache (`cache.provider: memcache`).
|
||
An account is compromised; an operator disables the user or the sessions are
|
||
revoked in the database, and the application calls `ClearUserCache(userID)`.
|
||
That returns an error and **removes nothing**. The attacker's cached
|
||
`UserContext` continues to authenticate every request for the full
|
||
`a.cacheTTL` — the database is never consulted again during that window
|
||
(`GetOrSet` short-circuits on a cache hit). Whether the operator even learns
|
||
this failed depends entirely on whether the caller checks the returned error;
|
||
`ClearUserCache` is also broken for a second reason — it ignores `userID` and
|
||
would have revoked every session in the process.
|
||
|
||
Note also that `ClearUserCache`'s pattern is *not* user-scoped, so even on Redis
|
||
and memory it is a global logout, not a per-user one. That direction is at least
|
||
fail-safe.
|
||
|
||
**Recommendation.** Revocation must not depend on a capability the provider may
|
||
not have. Options, in order of preference:
|
||
|
||
- Tag every session entry (`SetWithTags` with tags `auth:session`,
|
||
`auth:user:<id>`) and revoke with `DeleteByTag`, which all three providers
|
||
implement. This also makes `ClearUserCache` actually per-user.
|
||
- Keep a short `cacheTTL` (seconds, not minutes) so the revocation window is
|
||
bounded regardless.
|
||
- Make `DeleteByPattern`'s unsupported case loud: have `pkg/security` refuse to
|
||
start, or fall back to `Clear`, when the configured provider cannot revoke.
|
||
- At minimum, log at error level when a revocation call fails.
|
||
|
||
---
|
||
|
||
### 3. High — `defaultCache` is an unsynchronized lazy singleton on the request path
|
||
|
||
`pkg/cache/cache.go:9-62`:
|
||
|
||
```go
|
||
var (
|
||
defaultCache *Cache
|
||
)
|
||
|
||
func Initialize(provider Provider) {
|
||
defaultCache = NewCache(provider)
|
||
}
|
||
|
||
func UseMemory(opts *Options) error {
|
||
provider := NewMemoryProvider(opts)
|
||
defaultCache = NewCache(provider)
|
||
return nil
|
||
}
|
||
// ... UseRedis (:32), UseMemcache (:42) likewise
|
||
|
||
func GetDefaultCache() *Cache {
|
||
if defaultCache == nil {
|
||
_ = UseMemory(&Options{
|
||
DefaultTTL: 5 * time.Minute,
|
||
MaxSize: 10000,
|
||
})
|
||
}
|
||
return defaultCache
|
||
}
|
||
|
||
func SetDefaultCache(cache *Cache) {
|
||
defaultCache = cache
|
||
}
|
||
```
|
||
|
||
Six functions write `defaultCache` and `GetDefaultCache` both reads and writes
|
||
it, with no mutex, no `sync.Once` and no `atomic.Pointer`. `GetDefaultCache` is
|
||
called **from HTTP request handlers**: `pkg/restheadspec/handler.go:833`
|
||
(`cache.GetDefaultCache().Get(ctx, cacheKey, cachedTotalData)`),
|
||
`pkg/restheadspec/cache_helpers.go:109` and `:118`, and the equivalent
|
||
`pkg/resolvespec` paths.
|
||
|
||
Nothing in `pkg/` ever calls `Initialize` or `Use*`, so in a default deployment
|
||
**the first traffic to arrive is what initializes the cache**, concurrently.
|
||
|
||
**Failure scenario.** Two requests arrive simultaneously on a cold process.
|
||
Both observe `defaultCache == nil`, both run `UseMemory`, each constructing its
|
||
own `MemoryProvider`. One assignment wins. Request A writes its query total into
|
||
the provider that loses and is immediately garbage — so the entry is
|
||
unreachable, the cache reports a permanent miss for it, and the count is
|
||
recomputed from the database on every subsequent request. Worse, if this races
|
||
with an application's explicit `UseRedis` during startup, the Redis provider can
|
||
be clobbered by the lazy memory provider (or vice versa) and **the process
|
||
silently runs on the wrong backend**, which for `pkg/security` means session
|
||
cache entries that no other process shares and that `ClearCache` on another
|
||
instance can never reach.
|
||
|
||
This is also a genuine data race on the pointer word: unsynchronized
|
||
read/write of `defaultCache`, which `go test -race` would report immediately.
|
||
See `_CROSS-CUTTING.audit.md` — `-race` is never run in this repo, and
|
||
`pkg/cache` is not in the tested package set.
|
||
|
||
Note the same pattern exists in `pkg/security/providers.go:145`
|
||
(`cacheInstance = cache.GetDefaultCache()`), which at least happens at
|
||
construction time.
|
||
|
||
**Recommendation.** Guard the global with `sync.RWMutex` or store it in an
|
||
`atomic.Pointer[Cache]`, and make the lazy default a `sync.Once`:
|
||
|
||
```go
|
||
var (
|
||
defaultCache atomic.Pointer[Cache]
|
||
defaultOnce sync.Once
|
||
)
|
||
|
||
func GetDefaultCache() *Cache {
|
||
if c := defaultCache.Load(); c != nil {
|
||
return c
|
||
}
|
||
defaultOnce.Do(func() {
|
||
defaultCache.CompareAndSwap(nil, NewCache(NewMemoryProvider(&Options{
|
||
DefaultTTL: 5 * time.Minute, MaxSize: 10000,
|
||
})))
|
||
})
|
||
return defaultCache.Load()
|
||
}
|
||
```
|
||
|
||
Also: every replacement path drops the previous provider **without closing it**,
|
||
so `UseRedis` after a lazy `UseMemory` (or two `UseRedis` calls) leaks the old
|
||
provider's connection pool and, for Redis, its background goroutines. Close the
|
||
old provider on swap.
|
||
|
||
---
|
||
|
||
### 4. High — `tagToKeys` is unbounded; `MaxSize` bounds only `items`
|
||
|
||
`MemoryProvider` holds two maps (`provider_memory.go:30-37`):
|
||
|
||
```go
|
||
type MemoryProvider struct {
|
||
mu sync.RWMutex
|
||
items map[string]*memoryItem
|
||
tagToKeys map[string]map[string]struct{} // tag -> set of keys
|
||
options *Options
|
||
...
|
||
}
|
||
```
|
||
|
||
`MaxSize` is checked only against `len(m.items)` (`:105`, `:135`). Six paths
|
||
remove entries from `items` **without** removing them from `tagToKeys`:
|
||
|
||
| Path | Line | Cleans `tagToKeys`? |
|
||
|---|---|---|
|
||
| `Get` — expired-item delete | `:70` | no |
|
||
| `Set` — overwrites a tagged key | `:111` | no (and drops `Tags`, so the entry becomes unreachable for cleanup) |
|
||
| `evictOne` — expired scan | `:313` | no |
|
||
| `evictOne` — LRU victim | `:324` | no |
|
||
| `DeleteByPattern` | `:242` | no |
|
||
| `Clear` | `:254` | no — `m.tagToKeys` is never reset |
|
||
| `CleanExpired` | `:336` | no |
|
||
|
||
Only `Delete` (`:178-187`) and `SetWithTags` (`:141-151`) maintain it, and
|
||
`DeleteByTag` (`:226`) drops one whole tag.
|
||
|
||
Tags come from `pkg/restheadspec/cache_helpers.go:99-105`:
|
||
|
||
```go
|
||
func buildCacheTags(schema, tableName string) []string {
|
||
return []string{
|
||
fmt.Sprintf("schema:%s", strings.ToLower(schema)),
|
||
fmt.Sprintf("table:%s", strings.ToLower(tableName)),
|
||
}
|
||
}
|
||
```
|
||
|
||
and keys from `buildExtendedQueryCacheKey` (`:43-85`) — a SHA-256 of the full
|
||
query shape, including filters, sort, `customWhere`, `customOr`, `customJoin`,
|
||
expand and cursors.
|
||
|
||
**Failure scenario.** An attacker issues `GET /api/public/orders?...` in a loop,
|
||
varying one filter value each time. Every request produces a distinct SHA-256
|
||
key, `setQueryTotalCache` stores it under the tags `schema:public` and
|
||
`table:orders`, and `tagToKeys["table:orders"][key]` gains a member. Once
|
||
`items` reaches `MaxSize` (10 000 by default), `evictOne` starts discarding
|
||
items — but **never** the corresponding `tagToKeys` members. `items` stays
|
||
capped at 10 000; `tagToKeys["table:orders"]` grows by one 64-character key per
|
||
request, forever. At roughly 100 bytes per map entry, a few million requests —
|
||
easily reachable at modest rate — costs hundreds of megabytes of heap that
|
||
nothing will ever reclaim, because `Clear()` does not reset the map and
|
||
`CleanExpired` does not touch it. This is a **memory-exhaustion DoS driven
|
||
purely by query-string variation**, and it is cheap for the attacker: the
|
||
expensive part (the actual count query) can be avoided by hitting a table whose
|
||
count is trivial.
|
||
|
||
The leak also breaks invalidation correctness: `DeleteByTag` iterates a key set
|
||
full of keys that no longer exist, and a plain `Set` over a previously tagged
|
||
key leaves that key in the tag index while clearing its `Tags` — so the item can
|
||
be deleted by a tag it no longer claims.
|
||
|
||
**Recommendation.** Factor tag maintenance into a single private helper and call
|
||
it from every removal path:
|
||
|
||
```go
|
||
// caller must hold m.mu for writing
|
||
func (m *MemoryProvider) removeLocked(key string) {
|
||
if item, ok := m.items[key]; ok {
|
||
for _, tag := range item.Tags {
|
||
if ks := m.tagToKeys[tag]; ks != nil {
|
||
delete(ks, key)
|
||
if len(ks) == 0 {
|
||
delete(m.tagToKeys, tag)
|
||
}
|
||
}
|
||
}
|
||
}
|
||
delete(m.items, key)
|
||
}
|
||
```
|
||
|
||
Use it in `Get`'s expired path, `Set` (before overwrite), `evictOne`,
|
||
`DeleteByPattern` and `CleanExpired`; reset `m.tagToKeys` in `Clear`; and
|
||
account `len(m.tagToKeys)` (or total members) against an explicit bound.
|
||
Independently, cap the number of distinct tags and the members per tag.
|
||
|
||
---
|
||
|
||
### 5. High — `MemoryProvider.Get` takes the write lock on every hit
|
||
|
||
`provider_memory.go:56-88`:
|
||
|
||
```go
|
||
func (m *MemoryProvider) Get(ctx context.Context, key string) ([]byte, bool) {
|
||
// First try with read lock for fast path
|
||
m.mu.RLock()
|
||
item, exists := m.items[key]
|
||
...
|
||
value := item.Value
|
||
m.mu.RUnlock()
|
||
|
||
// Update access tracking with write lock
|
||
m.mu.Lock()
|
||
item.LastAccess = time.Now()
|
||
item.HitCount++
|
||
m.mu.Unlock()
|
||
|
||
m.hits.Add(1)
|
||
return value, true
|
||
}
|
||
```
|
||
|
||
The comment promises a read-lock fast path, but every **successful** lookup ends
|
||
in an exclusive lock to bump two bookkeeping fields. The `RWMutex` therefore
|
||
provides no read concurrency at all on the hot path, and Go's `RWMutex` blocks
|
||
*new* readers once a writer is waiting — so a burst of concurrent hits
|
||
degenerates into a fully serialized queue with two lock handoffs per operation.
|
||
|
||
**Failure scenario.** The session cache is the default `MemoryProvider`. Under
|
||
concurrent load, every authenticated request performs
|
||
`GetOrSet` → `Get` → cache hit → exclusive lock. With hundreds of in-flight
|
||
requests the mutex becomes the throughput ceiling for the entire API, and
|
||
because the write lock is taken *after* the read lock is released, each hit pays
|
||
two full lock acquisitions. Adding `HitCount` to a per-item `atomic.Int64`
|
||
would make this free; as written, the cache that exists to reduce latency is the
|
||
serialization point.
|
||
|
||
A secondary defect: between `RUnlock` at `:78` and `Lock` at `:81` the item may
|
||
have been deleted or replaced, so the code can mutate an orphaned struct. Harmless
|
||
but confirms the bookkeeping does not need the lock.
|
||
|
||
**Recommendation.** Make the counters lock-free and drop the write lock:
|
||
|
||
```go
|
||
type memoryItem struct {
|
||
Value []byte
|
||
Expiration time.Time
|
||
lastAccess atomic.Int64 // unix nanos
|
||
hitCount atomic.Int64
|
||
Tags []string
|
||
}
|
||
```
|
||
|
||
Then the whole `Get` runs under `RLock`. If exact LRU ordering matters, consider
|
||
an approximate clock (update `lastAccess` only if it is more than a second
|
||
stale) or a sharded map to cut contention.
|
||
|
||
---
|
||
|
||
### 6. High — `DeleteByPattern` has three incompatible pattern languages
|
||
|
||
The interface (`provider.go:29-31`) says only "Pattern syntax depends on the
|
||
provider implementation", and the three implementations diverge completely:
|
||
|
||
| Provider | Line | Semantics |
|
||
|---|---|---|
|
||
| memory | `provider_memory.go:235-244` | `regexp.Compile` + **unanchored** `MatchString` |
|
||
| redis | `provider_redis.go:197` | `SCAN MATCH` — Redis glob |
|
||
| memcache | `provider_memcache.go:252-254` | always an error |
|
||
|
||
```go
|
||
// memory
|
||
re, err := regexp.Compile(pattern)
|
||
if err != nil {
|
||
return fmt.Errorf("invalid pattern: %w", err)
|
||
}
|
||
for key := range m.items {
|
||
if re.MatchString(key) {
|
||
delete(m.items, key)
|
||
}
|
||
}
|
||
```
|
||
|
||
The one caller, `pkg/security/providers.go:470`, passes `"auth:session:*"` —
|
||
a Redis glob. Interpreted as a Go regexp that is `auth:session` followed by zero
|
||
or more `:`, matched unanchored, so it happens to match the intended keys (and
|
||
any key merely *containing* `auth:session`). It works by coincidence, not design.
|
||
|
||
**Failure scenario.** Two ways this bites:
|
||
|
||
- **Over-deletion.** Because matching is unanchored, a glob like `user:*` becomes
|
||
the regexp `user:*` = `user` + zero-or-more colons, which matches *any* key
|
||
containing `user` — including `auth:session:<token>` if a token happens to
|
||
contain that substring. Conversely a caller who writes a glob such as
|
||
`*` gets `regexp.Compile("*")` → `error parsing regexp: missing argument to
|
||
repetition operator`, i.e. a silent no-op invalidation where the author
|
||
expected a full flush. Stale authorization data continues to be served.
|
||
- **Attacker-influenced regexp.** `regexp.Compile` runs on the caller's string
|
||
**while holding the write lock** (`:232-238`), so if any future caller derives
|
||
a pattern from request input, an attacker both controls the compiled program
|
||
and blocks every other cache operation for its duration. Go's RE2 has no
|
||
catastrophic backtracking, but compilation of a large pattern is not free and
|
||
the lock is held across it.
|
||
|
||
Redis's side has its own cost: `r.client.Scan(ctx, 0, pattern, 0)` with
|
||
`count = 0` leaves the server at its default `COUNT 10`, so revoking sessions
|
||
walks the entire keyspace in ~10-key increments — thousands of round trips on a
|
||
large DB, executed synchronously inside `ClearCache`.
|
||
|
||
**Recommendation.** Define one pattern language in the interface — a glob is the
|
||
right choice, since it is the one Redis supports natively — and implement it for
|
||
memory with `path.Match` (or an explicit anchored translation to regexp),
|
||
compiled **before** taking the lock. Reject patterns the provider cannot honour
|
||
with a typed `ErrUnsupported` so callers can branch. Pass a sensible `COUNT`
|
||
(e.g. 500) to `SCAN`. Better still, replace pattern deletion with tag deletion
|
||
at the one call site (see finding 2).
|
||
|
||
---
|
||
|
||
### 7. High — the session token is embedded in an error string
|
||
|
||
`cache_manager.go:22-43`:
|
||
|
||
```go
|
||
func (c *Cache) Get(ctx context.Context, key string, dest interface{}) error {
|
||
data, exists := c.provider.Get(ctx, key)
|
||
if !exists {
|
||
return fmt.Errorf("key not found: %s", key)
|
||
}
|
||
...
|
||
}
|
||
|
||
func (c *Cache) GetBytes(ctx context.Context, key string) ([]byte, error) {
|
||
data, exists := c.provider.Get(ctx, key)
|
||
if !exists {
|
||
return nil, fmt.Errorf("key not found: %s", key)
|
||
}
|
||
return data, nil
|
||
}
|
||
```
|
||
|
||
The key for the session cache is `"auth:session:" + token` — the raw bearer
|
||
credential. Every cache miss therefore allocates an error whose text contains a
|
||
live secret. `GetOrSet` discards it, but this is a public API on a `*Cache` that
|
||
`pkg/security` holds directly, and `pkg/security/keystore_database.go:232` calls
|
||
`ks.cache.Get` on an API-key cache too.
|
||
|
||
**Failure scenario.** Any caller that does `logger.Error("cache lookup failed: %v", err)`
|
||
publishes the bearer token to the application log **and**, per
|
||
`audit/pkg/logger.audit.md` finding 2, forwards it verbatim to Sentry, where it
|
||
is retained by a third party with no scrubbing (`pkg/errortracking` has no
|
||
`BeforeSend` hook). A token in a log aggregator is a replayable credential for
|
||
the whole `cacheTTL` — longer, if the log outlives the session. Note that this
|
||
requires only one careless `%v` at a call site; the package is handing out the
|
||
loaded weapon.
|
||
|
||
**Recommendation.** Never interpolate cache keys into errors. Use a package
|
||
sentinel and let the caller decide what is safe to log:
|
||
|
||
```go
|
||
var ErrNotFound = errors.New("cache: key not found")
|
||
...
|
||
if !exists {
|
||
return ErrNotFound
|
||
}
|
||
```
|
||
|
||
Callers then use `errors.Is(err, cache.ErrNotFound)` instead of matching on
|
||
strings, which also fixes the miss/error conflation noted in finding 17. If a
|
||
key must appear in diagnostics, log a truncated hash of it. Separately,
|
||
`pkg/security` should key the cache on a SHA-256 of the token rather than the
|
||
token itself, exactly as `keystore_database.go` already does for API keys
|
||
(`keystoreCacheKey(hash)`, `:287`).
|
||
|
||
---
|
||
|
||
### 8. High — no single-flight: concurrent misses run N loaders
|
||
|
||
`GetOrSet` (`cache_manager.go:112`) and `Remember` (`:145`) both do
|
||
check → load → store with nothing serializing concurrent callers on the same
|
||
key.
|
||
|
||
**Failure scenario (auth).** A client opens 200 connections with the same fresh
|
||
session token. All 200 miss the cache, all 200 enter the loader, and all 200
|
||
execute the session stored procedure
|
||
(`SELECT p_success, p_error, p_user::text FROM <session_fn>($1, $2)`,
|
||
`pkg/security/providers.go:413-415`) concurrently. Each success then also spawns
|
||
`go a.updateSessionActivity(...)` (`:447`), i.e. 200 more goroutines each issuing
|
||
a database write. The cache provides no protection at all for the first
|
||
round-trip, and under sustained concurrency where request arrival outpaces query
|
||
latency it never catches up: throughput is bounded by the database, not the
|
||
cache. A single valid credential is enough to drive this — no privilege needed.
|
||
|
||
**Failure scenario (query totals).** The same shape applies to
|
||
`restheadspec/handler.go:833`: a burst of identical expensive `COUNT(*)` queries
|
||
all miss together and all hit the database.
|
||
|
||
This is the classic cache stampede / thundering herd, and it is the reason
|
||
single-flight exists.
|
||
|
||
**Recommendation.** Wrap the loader in `golang.org/x/sync/singleflight`, which is
|
||
already an indirect dependency of most Go service stacks:
|
||
|
||
```go
|
||
type Cache struct {
|
||
provider Provider
|
||
sf singleflight.Group
|
||
}
|
||
|
||
func (c *Cache) GetOrSet(ctx context.Context, key string, dest any, ttl time.Duration, loader func() (any, error)) error {
|
||
if err := c.Get(ctx, key, dest); err == nil {
|
||
return nil
|
||
}
|
||
v, err, _ := c.sf.Do(key, func() (any, error) {
|
||
// re-check under the flight, then load and store
|
||
...
|
||
})
|
||
...
|
||
}
|
||
```
|
||
|
||
Note the group must be keyed per-`Cache`, and `Forget` should be called on
|
||
loader error so a failure is not shared beyond the in-flight set. For the auth
|
||
path specifically, also bound `updateSessionActivity` — an unbounded `go` per
|
||
request is its own DoS vector (raised again in the `pkg/security` audit).
|
||
|
||
---
|
||
|
||
### 9. Medium — attacker-controlled keys violate memcache's key rules
|
||
|
||
gomemcache enforces the protocol limits (verified in
|
||
`gomemcache@v0.0.0-20260422231931-4d751bb6e37c/memcache.go:58-91`):
|
||
|
||
```go
|
||
// ErrMalformedKey is returned when an invalid key is used.
|
||
// Keys must be at maximum 250 bytes long and not
|
||
// contain whitespace or control characters.
|
||
ErrMalformedKey = errors.New("malformed: key is too long or contains invalid characters")
|
||
|
||
func legalKey(key string) bool {
|
||
if len(key) > 250 {
|
||
return false
|
||
}
|
||
...
|
||
}
|
||
```
|
||
|
||
The session cache key is `"auth:session:" + token` with `token` taken from the
|
||
`Authorization` header, so its length and byte content are chosen by the client.
|
||
|
||
**Failure scenario.** A deployment uses memcache and issues JWT session tokens,
|
||
which routinely exceed 238 bytes. `MemcacheProvider.Get` receives
|
||
`ErrMalformedKey` and — per finding 17 — reports it as a plain cache miss
|
||
(`provider_memcache.go:80-82`). The loader runs, the database validates the
|
||
session, and then `MemcacheProvider.Set` returns `ErrMalformedKey`, which
|
||
finding 1 converts into an authentication failure. **Every user with a long
|
||
token is permanently unable to authenticate**, and the logs show nothing but a
|
||
generic "failed to cache value". A client can also trigger this deliberately
|
||
with a token containing a space to probe the backend.
|
||
|
||
**Recommendation.** Hash keys inside the provider so key length and charset are
|
||
bounded regardless of caller input — e.g. `sha256` hex of the key when it
|
||
exceeds 200 bytes or contains illegal bytes, with a fixed prefix. And, as in
|
||
finding 7, `pkg/security` should hash the token before it ever becomes a key.
|
||
|
||
---
|
||
|
||
### 10. Medium — TOCTOU when deleting an expired item
|
||
|
||
`provider_memory.go:66-74`:
|
||
|
||
```go
|
||
if item.isExpired() {
|
||
m.mu.RUnlock()
|
||
// Upgrade to write lock to delete expired item
|
||
m.mu.Lock()
|
||
delete(m.items, key)
|
||
m.mu.Unlock()
|
||
m.misses.Add(1)
|
||
return nil, false
|
||
}
|
||
```
|
||
|
||
Go's `RWMutex` has no lock upgrade, so the read lock is genuinely released and
|
||
the write lock separately acquired. The `delete` is then **unconditional** — it
|
||
does not re-check that the entry still exists or is still the expired one.
|
||
|
||
**Failure scenario.** Goroutine A reads an expired entry for key `K` and drops
|
||
the read lock. Goroutine B takes the write lock and `Set`s a fresh value for
|
||
`K`. Goroutine A now takes the write lock and deletes B's fresh entry. B has no
|
||
idea; the value it believes it cached is gone, and the next reader recomputes
|
||
it. In the auth path this means a just-validated session is dropped immediately,
|
||
forcing another stored-procedure call — and under load this can repeat, since the
|
||
interleaving recurs whenever expiry and refresh coincide, which is exactly when
|
||
traffic for that key is highest.
|
||
|
||
**Recommendation.** Re-check under the write lock and delete only the same
|
||
entry:
|
||
|
||
```go
|
||
m.mu.Lock()
|
||
if cur, ok := m.items[key]; ok && cur == item {
|
||
m.removeLocked(key) // see finding 4
|
||
}
|
||
m.mu.Unlock()
|
||
```
|
||
|
||
---
|
||
|
||
### 11. Medium — `Close()` makes the provider panic on next use
|
||
|
||
`provider_memory.go:273-280`:
|
||
|
||
```go
|
||
func (m *MemoryProvider) Close() error {
|
||
m.mu.Lock()
|
||
defer m.mu.Unlock()
|
||
|
||
m.items = nil
|
||
return nil
|
||
}
|
||
```
|
||
|
||
Reads of a nil map are fine, but `Set`/`SetWithTags` do
|
||
`m.items[key] = &memoryItem{...}` (`:111`, `:154`), which on a nil map panics
|
||
with `assignment to entry in nil map`.
|
||
|
||
**Failure scenario.** A graceful-shutdown handler calls `cache.Close()` while
|
||
requests are still draining — the normal ordering, since `pkg/config` defaults
|
||
`servers.drain_timeout` to 25s. The next in-flight request reaches
|
||
`setQueryTotalCache` and the process panics **while holding `m.mu`**. There is
|
||
no `recover()` anywhere in `pkg/cache`, so this unwinds into whatever the caller
|
||
has; if the HTTP server's panic handler recovers it, the mutex is never
|
||
unlocked and **every subsequent cache operation blocks forever** — the process
|
||
is alive, serving, and permanently wedged on the first cache access. That is a
|
||
worse outcome than crashing.
|
||
|
||
**Recommendation.** Mark the provider closed instead of destroying the map, and
|
||
return an error from every method afterwards:
|
||
|
||
```go
|
||
type MemoryProvider struct {
|
||
...
|
||
closed bool
|
||
}
|
||
|
||
func (m *MemoryProvider) Close() error {
|
||
m.mu.Lock()
|
||
defer m.mu.Unlock()
|
||
m.closed = true
|
||
m.items = make(map[string]*memoryItem)
|
||
m.tagToKeys = make(map[string]map[string]struct{})
|
||
return nil
|
||
}
|
||
```
|
||
|
||
with `if m.closed { return ErrClosed }` at the top of each mutating method.
|
||
More generally, the package should use `defer logger.CatchPanicCallback(...)` at
|
||
its exported boundaries so a panic inside a lock is reported rather than
|
||
silently converted into a deadlock — but note `audit/pkg/logger.audit.md`
|
||
finding 4 on `CatchPanic` swallowing unconditionally; here the panic should be
|
||
logged and **re-raised** or converted to an error, not absorbed.
|
||
|
||
---
|
||
|
||
### 12. Medium — the memcache tag index is a lost-update machine
|
||
|
||
`provider_memcache.go:136-170` maintains, for each tag, a JSON array of keys:
|
||
|
||
```go
|
||
for _, tag := range tags {
|
||
tagKey := fmt.Sprintf("cache:tag:%s", tag)
|
||
|
||
// Get existing keys for this tag
|
||
var keys []string
|
||
if item, err := m.client.Get(tagKey); err == nil {
|
||
_ = json.Unmarshal(item.Value, &keys)
|
||
}
|
||
|
||
// Add current key if not already present
|
||
found := false
|
||
for _, k := range keys { ... }
|
||
if !found {
|
||
keys = append(keys, key)
|
||
}
|
||
|
||
keysData, err := json.Marshal(keys)
|
||
if err != nil {
|
||
continue
|
||
}
|
||
|
||
tagItem := &memcache.Item{
|
||
Key: tagKey,
|
||
Value: keysData,
|
||
Expiration: expiration + 3600, // Give tag lists longer TTL
|
||
}
|
||
_ = m.client.Set(tagItem)
|
||
}
|
||
```
|
||
|
||
Four distinct defects in fifteen lines:
|
||
|
||
1. **Non-atomic read-modify-write.** `Get` then `Set` with no CAS, across a
|
||
network, for a value that every concurrent writer of the same tag touches.
|
||
2. **Unbounded growth.** The list only ever grows within its TTL, and every
|
||
writer re-reads and re-writes the whole array. `Delete` (`:190-200`) rewrites
|
||
it too.
|
||
3. **1 MB item limit.** Once the array exceeds memcached's default item size the
|
||
`Set` fails — and the error is discarded by `_ =`.
|
||
4. **`expiration + 3600` can cross memcached's 30-day boundary.** The protocol
|
||
treats an expiry value above 2 592 000 as an **absolute Unix timestamp**. A
|
||
caller passing a 30-day TTL yields `2592000 + 3600 = 2595600`, which
|
||
memcached reads as 1970-01-31 — already past, so the tag list is dead on
|
||
arrival. `int32(ttl.Seconds())` at `:108` also truncates for very large TTLs
|
||
and, for a negative TTL, expires the item immediately.
|
||
|
||
**Failure scenario.** Two requests cache different query totals for the same
|
||
table concurrently. Both read `cache:tag:table:orders` and see `[k1]`; one
|
||
writes `[k1,k2]`, the other writes `[k1,k3]`. The second write wins and `k2` is
|
||
**no longer in the tag index**. A later `POST` to that table calls
|
||
`invalidateCacheForTags` → `DeleteByTag("table:orders")`, which deletes `k1` and
|
||
`k3` but not `k2`. `k2` continues to serve a **stale row count from before the
|
||
write** for its full TTL. Under concurrency this is not an edge case — it is the
|
||
normal outcome, and lost updates accumulate. When the list crosses 1 MB the
|
||
failure becomes total: no further keys are indexed and nothing is logged.
|
||
|
||
Additionally, the prefixes `cache:tag:` and `cache:tags:` (`:128`, `:138`,
|
||
`:179`, `:185`, `:219`, `:239`) share the key namespace with ordinary cache
|
||
keys. No consumer currently writes keys under that prefix — the restheadspec
|
||
keys are `query_total:<sha256>` and the security keys `auth:session:*` — but
|
||
nothing enforces it, and a consumer that ever allows an attacker-derived key
|
||
beginning `cache:tag:` could forge or destroy the invalidation index (mass
|
||
invalidation → database load, or suppressed invalidation → stale authorization
|
||
data served).
|
||
|
||
**Recommendation.** Use `CompareAndSwap` (gomemcache exposes it) in a bounded
|
||
retry loop, cap the list length, honour the 30-day rule, and stop discarding
|
||
errors:
|
||
|
||
```go
|
||
func memcacheExpiry(ttl time.Duration) int32 {
|
||
secs := int64(ttl.Seconds())
|
||
if secs < 0 { secs = 0 }
|
||
if secs > 2592000 { // >30d must be an absolute timestamp
|
||
return int32(time.Now().Add(ttl).Unix())
|
||
}
|
||
return int32(secs)
|
||
}
|
||
```
|
||
|
||
Given how weak tag support is on memcache, the honest alternative is to return
|
||
`ErrUnsupported` from `SetWithTags`/`DeleteByTag` and force callers to pick a
|
||
provider that can do it, rather than offering invalidation that silently misses
|
||
keys. Namespace the index keys under a prefix that ordinary keys cannot reach
|
||
(e.g. by prefixing all user keys with `k:`).
|
||
|
||
---
|
||
|
||
### 13. Medium — `Clear()` flushes the entire shared server
|
||
|
||
```go
|
||
// provider_memcache.go:257-259
|
||
func (m *MemcacheProvider) Clear(ctx context.Context) error {
|
||
return m.client.FlushAll()
|
||
}
|
||
|
||
// provider_redis.go:228-230
|
||
func (r *RedisProvider) Clear(ctx context.Context) error {
|
||
return r.client.FlushDB(ctx).Err()
|
||
}
|
||
```
|
||
|
||
Neither is scoped to this application's keys. `FlushAll` wipes every key on
|
||
every configured memcached server; `FlushDB` wipes the whole logical Redis DB —
|
||
and `RedisConfig.DB` defaults to `0`, which is also where `pkg/config` puts the
|
||
event broker (`event_broker.redis.db: 0`, `pkg/config/manager.go:265`) and its
|
||
`resolvespec:events` stream.
|
||
|
||
**Failure scenario.** An operator or an admin endpoint calls `cache.Clear()` to
|
||
drop stale entries. On Redis with the default config this also deletes the event
|
||
broker's stream and consumer-group state, so queued events are lost and
|
||
consumers fail; on memcached it evicts every other tenant sharing that server.
|
||
There is no confirmation, no scoping, and the method is one call away from any
|
||
consumer holding the `*Cache`.
|
||
|
||
**Recommendation.** Implement `Clear` as a scoped delete over the application's
|
||
own key prefix (`SCAN`+`DEL` for Redis), require an explicitly opted-in
|
||
"destructive flush" flag to use `FlushDB`/`FlushAll`, and document that the
|
||
cache must not share a Redis DB with the event broker. Consider adding a
|
||
mandatory `KeyPrefix` to `Options` so scoping is always possible.
|
||
|
||
---
|
||
|
||
### 14. Medium — `MemcacheProvider.Close()` is a no-op based on a false comment
|
||
|
||
`provider_memcache.go:267-271`:
|
||
|
||
```go
|
||
func (m *MemcacheProvider) Close() error {
|
||
// Memcache client doesn't have a close method
|
||
return nil
|
||
}
|
||
```
|
||
|
||
This is factually wrong for the pinned dependency: gomemcache
|
||
`v0.0.0-20260422231931-4d751bb6e37c` exposes `func (c *Client) Close() error` at
|
||
`memcache.go:836`. Its documented behaviour is to close all currently-open idle
|
||
connections (the client stays usable afterwards), which is exactly what a
|
||
provider `Close()` should be releasing.
|
||
|
||
**Failure scenario.** Every provider replacement (finding 3) and every
|
||
`cache.Close()` leaves the memcache client's idle connections — up to
|
||
`MaxIdleConns` per server — established. In a process that reconfigures the
|
||
cache, or in tests that construct providers repeatedly, file descriptors
|
||
accumulate until `accept`/`dial` starts failing with `too many open files`,
|
||
which manifests as unrelated failures elsewhere in the process.
|
||
|
||
**Recommendation.** `return m.client.Close()`. Also add `Close` handling to the
|
||
swap paths in `cache.go` per finding 3.
|
||
|
||
---
|
||
|
||
### 15. Medium — no TLS option for Redis or memcache
|
||
|
||
`RedisConfig` (`provider_redis.go:17-36`) has `Host`, `Port`, `Password`, `DB`,
|
||
`PoolSize`, `Options` — and nothing else. `redis.Options` supports `TLSConfig`,
|
||
but the constructor never sets it:
|
||
|
||
```go
|
||
client := redis.NewClient(&redis.Options{
|
||
Addr: fmt.Sprintf("%s:%d", config.Host, config.Port),
|
||
Password: config.Password,
|
||
DB: config.DB,
|
||
PoolSize: config.PoolSize,
|
||
})
|
||
```
|
||
|
||
`MemcacheConfig` (`:18-31`) likewise has no transport security, and the
|
||
memcached protocol has no in-band auth here at all.
|
||
|
||
**Failure scenario.** The cached values are `UserContext` objects — identity and
|
||
authorization data — and the `AUTH` password is sent in cleartext on the first
|
||
command of every new connection. Anything able to observe the path between the
|
||
service and Redis (a shared VPC, a misconfigured security group, a compromised
|
||
sidecar, a managed Redis reached over the public internet) can read session
|
||
contents and steal the Redis password, then write forged `auth:session:*` entries
|
||
directly. Writing a crafted `UserContext` into the cache is a **complete
|
||
authentication bypass**: `GetOrSet` returns it on a hit and never consults the
|
||
database.
|
||
|
||
This is the same theme as `pkg/config`'s `sslmode: disable` default
|
||
(`pkg/config/manager.go:242`) — see `audit/pkg/config.audit.md` finding 3.
|
||
|
||
**Recommendation.** Add TLS configuration to both configs and plumb it through:
|
||
|
||
```go
|
||
type RedisConfig struct {
|
||
...
|
||
TLS bool
|
||
TLSSkipVerify bool // must default false
|
||
TLSCACertFile string
|
||
}
|
||
```
|
||
|
||
and set `redis.Options.TLSConfig` accordingly. Since the cache holds
|
||
authentication material, treat encrypted transport as the default and require an
|
||
explicit opt-out. For memcache, prefer Redis for this workload or terminate TLS
|
||
with a local proxy (stunnel/envoy) and document it.
|
||
|
||
---
|
||
|
||
### 16. Medium — `Remember` returns a different type on hit vs miss
|
||
|
||
`cache_manager.go:145-167`:
|
||
|
||
```go
|
||
func (c *Cache) Remember(ctx context.Context, key string, ttl time.Duration, loader func() (interface{}, error)) (interface{}, error) {
|
||
data, err := c.GetBytes(ctx, key)
|
||
if err == nil {
|
||
var result interface{}
|
||
if err := json.Unmarshal(data, &result); err == nil {
|
||
return result, nil // <-- map[string]interface{} / float64 / ...
|
||
}
|
||
}
|
||
|
||
value, err := loader()
|
||
...
|
||
return value, nil // <-- whatever the loader returned
|
||
}
|
||
```
|
||
|
||
On a hit the value comes back as generic JSON (`map[string]interface{}`,
|
||
`[]interface{}`, `float64`, `string`); on a miss it is the loader's concrete Go
|
||
type.
|
||
|
||
**Failure scenario.** A caller writes the natural thing:
|
||
|
||
```go
|
||
v, err := c.Remember(ctx, key, ttl, func() (any, error) { return loadUser(id) })
|
||
u := v.(*User) // panics on every cache hit
|
||
```
|
||
|
||
This passes every test run against a cold cache and panics in production as soon
|
||
as the cache warms — the worst possible failure timing. `Remember` has no
|
||
callers in `pkg/` today, so this is latent, but it is a trap laid for the next
|
||
consumer and there is no `recover()` in the package to contain it.
|
||
|
||
**Recommendation.** Either delete `Remember` in favour of `GetOrSet` (which
|
||
takes a typed `dest`), or give it the same contract:
|
||
|
||
```go
|
||
func (c *Cache) Remember(ctx context.Context, key string, dest any, ttl time.Duration, loader func() (any, error)) error
|
||
```
|
||
|
||
If the generic-return shape must stay, document it loudly and name it
|
||
`RememberAny`.
|
||
|
||
---
|
||
|
||
### 17. Medium — `Get`/`Exists` swallow every backend error as a cache miss
|
||
|
||
```go
|
||
// provider_redis.go:86-95
|
||
val, err := r.client.Get(ctx, key).Bytes()
|
||
if err == redis.Nil {
|
||
return nil, false
|
||
}
|
||
if err != nil {
|
||
return nil, false // network error, WRONGTYPE, auth failure, timeout...
|
||
}
|
||
|
||
// provider_memcache.go:75-84 — same shape
|
||
// provider_memcache.go:262-265
|
||
func (m *MemcacheProvider) Exists(ctx context.Context, key string) bool {
|
||
_, err := m.client.Get(key)
|
||
return err == nil
|
||
}
|
||
```
|
||
|
||
The `Provider.Get` signature `([]byte, bool)` cannot express "I don't know", so
|
||
every failure is indistinguishable from an absent key. Nothing is logged —
|
||
`pkg/cache` imports no logger.
|
||
|
||
**Failure scenario.** Redis develops packet loss, or memcached is restarted, or
|
||
`maxmemory` is hit. Every lookup reports a miss, so every request falls through
|
||
to the database: the count query at `restheadspec/handler.go:857` and the session
|
||
stored procedure at `providers.go:413`. Load multiplies by the cache hit ratio
|
||
in an instant — typically 10–100× — and the database becomes the next thing to
|
||
fall over. Meanwhile the metrics say "cache miss", the logs say nothing, and the
|
||
`Stats` endpoint reports a plausible-looking miss count, so the actual cause is
|
||
invisible during the incident. Combined with finding 1, the subsequent `Set`
|
||
failure also rejects every request, so the symptom presented to operators is
|
||
"authentication is broken" with no mention of Redis.
|
||
|
||
**Recommendation.** Widen the interface to `Get(ctx, key) ([]byte, bool, error)`
|
||
(or keep the two-value form and add `GetE`), and at minimum log at warn level
|
||
with a rate limit inside each provider:
|
||
|
||
```go
|
||
if err != nil && err != redis.Nil {
|
||
logger.Warn("cache: redis GET failed for prefix %s: %v", ctx, keyPrefix(key), err)
|
||
return nil, false
|
||
}
|
||
```
|
||
|
||
Note `keyPrefix` rather than `key`, per finding 7. Feed a
|
||
`cache_errors_total{provider}` counter into `pkg/metrics` so a degraded backend
|
||
is alertable — the `Provider` interface there already has `RecordCacheHit`/
|
||
`RecordCacheMiss` but no error counter.
|
||
|
||
---
|
||
|
||
### 18. Low — `evictOne` is an O(n) scan under the write lock
|
||
|
||
`provider_memory.go:306-326`:
|
||
|
||
```go
|
||
func (m *MemoryProvider) evictOne() {
|
||
var oldestKey string
|
||
var oldestTime time.Time
|
||
|
||
for key, item := range m.items {
|
||
if item.isExpired() {
|
||
delete(m.items, key)
|
||
return
|
||
}
|
||
if oldestKey == "" || item.LastAccess.Before(oldestTime) {
|
||
oldestKey = key
|
||
oldestTime = item.LastAccess
|
||
}
|
||
}
|
||
if oldestKey != "" {
|
||
delete(m.items, oldestKey)
|
||
}
|
||
}
|
||
```
|
||
|
||
Called from `Set` (`:107`) and `SetWithTags` (`:137`) whenever the cache is at
|
||
capacity — i.e. on **every insertion** in steady state — while the write lock is
|
||
held. With the default `MaxSize: 10000` that is a 10 000-entry map walk plus
|
||
10 000 `time.Time` comparisons per write, blocking all readers (which, per
|
||
finding 5, includes every cache hit).
|
||
|
||
`Stats` (`:283-304`) has the same O(n) shape under `RLock`, and its comment
|
||
"Clean expired items first" is wrong — it holds a **read** lock and only counts.
|
||
|
||
**Recommendation.** Use a proper LRU (an intrusive doubly-linked list beside the
|
||
map, as `hashicorp/golang-lru` or `container/list` gives you) for O(1) eviction,
|
||
or sample k random entries and evict the oldest of those (Redis's approach) if
|
||
approximate LRU is acceptable. Maintain a counter for `Stats` instead of
|
||
walking. Fix the misleading comment.
|
||
|
||
---
|
||
|
||
### 19. Low — `CleanExpired` is dead code; there is no janitor
|
||
|
||
`provider_memory.go:329` `CleanExpired` has no callers anywhere in the
|
||
repository (verified by grep — the only hits are its own declaration and
|
||
comment). No goroutine sweeps expirations.
|
||
|
||
**Failure scenario.** Expired entries are reclaimed only when someone looks them
|
||
up (`Get`, `:66`) or when `evictOne` happens to walk past one. A workload whose
|
||
keys are written once and never re-read — which is exactly the
|
||
`query_total:<sha256>` pattern, since a distinct filter combination is usually
|
||
requested once — retains every expired entry until `MaxSize` forces eviction.
|
||
The cache therefore sits permanently at its maximum footprint holding mostly
|
||
dead data, and the LRU scan in finding 18 walks those dead entries on every
|
||
insertion. With `MaxSize <= 0` (finding 22) nothing reclaims them at all.
|
||
|
||
**Recommendation.** Start a janitor goroutine from `NewMemoryProvider` with a
|
||
configurable interval, and stop it in `Close()`:
|
||
|
||
```go
|
||
func NewMemoryProvider(opts *Options) *MemoryProvider {
|
||
m := &MemoryProvider{...; done: make(chan struct{})}
|
||
go m.janitor(opts.CleanupInterval) // default e.g. 1 minute
|
||
return m
|
||
}
|
||
```
|
||
|
||
Guard the goroutine with `defer logger.CatchPanicCallback("cache.janitor", ...)`
|
||
so a panic in the sweep is reported rather than killing the process silently.
|
||
|
||
---
|
||
|
||
### 20. Low — `RedisProvider.Stats` returns the raw `INFO` output
|
||
|
||
`provider_redis.go:247-268`:
|
||
|
||
```go
|
||
info, err := r.client.Info(ctx, "stats", "keyspace").Result()
|
||
...
|
||
stats := &CacheStats{
|
||
Keys: dbSize,
|
||
ProviderType: "redis",
|
||
ProviderStats: map[string]any{
|
||
"info": info,
|
||
},
|
||
}
|
||
```
|
||
|
||
`CacheStats.ProviderStats` is `json:"provider_stats,omitempty"` — i.e. designed
|
||
to be serialized. The `stats` and `keyspace` sections include the Redis version,
|
||
uptime, connected-client counts, keyspace hit/miss totals, eviction and
|
||
expiration counters, and per-database key counts.
|
||
|
||
**Failure scenario.** `cache.GetStats` (`cache.go:65`) has no callers today, but
|
||
it is an obvious thing to wire to a `/health` or `/admin/stats` endpoint. Doing
|
||
so exposes infrastructure fingerprinting to any client that reaches it, and the
|
||
keyspace counters leak activity volume. It is an information-disclosure
|
||
primitive waiting for a route.
|
||
|
||
**Recommendation.** Parse `INFO` into a small allowlisted set of numeric fields
|
||
(`keyspace_hits`, `keyspace_misses`, `evicted_keys`, `expired_keys`) and
|
||
populate `CacheStats.Hits`/`Misses` from them rather than passing the blob
|
||
through. If the raw text is wanted for debugging, gate it behind an explicit
|
||
debug flag and never include it in a response body.
|
||
|
||
---
|
||
|
||
### 21. Low — the memcache provider ignores `ctx` entirely
|
||
|
||
Every method on `MemcacheProvider` accepts `ctx context.Context` and none uses
|
||
it (`:75`, `:87`, `:103`, `:177`, `:218`, `:252`, `:257`, `:262`, `:275`). The
|
||
only timeout is the client-wide `config.Timeout` (default 1s, `:49-51`).
|
||
|
||
**Failure scenario.** A client disconnects or the request deadline expires; the
|
||
handler's context is cancelled, but the cache call continues to completion. In
|
||
`SetWithTags` that is one `Set` plus, per tag, a `Get` and a `Set` — so a
|
||
two-tag write is five sequential round trips, each able to consume the full 1s
|
||
client timeout, all after the caller has given up. Under load-shedding
|
||
conditions the server keeps doing work for requests nobody is waiting for, which
|
||
is precisely when it can least afford to.
|
||
|
||
**Recommendation.** gomemcache's API is context-free, so either wrap each call
|
||
with a `select` on `ctx.Done()` and a goroutine, or switch to a
|
||
context-aware client. At minimum, check `ctx.Err()` at the top of each method
|
||
and return early, and document that `config.Timeout` is the real bound.
|
||
|
||
---
|
||
|
||
### 22. Low — `MaxSize <= 0` silently disables eviction
|
||
|
||
`provider_memory.go:105` and `:135` both guard with `if m.options.MaxSize > 0 && ...`.
|
||
A caller who constructs `&Options{DefaultTTL: time.Minute}` — as
|
||
`NewRedisProvider` (`:59`) and `NewMemcacheProvider` (`:54`) do for their own
|
||
defaults, and as any hand-written config easily does — gets an **unbounded**
|
||
in-memory cache. Combined with findings 4 and 19, memory then grows with
|
||
attacker-controlled query variation until the process is OOM-killed.
|
||
|
||
**Recommendation.** Treat `MaxSize <= 0` as "use the default" (10 000) rather
|
||
than "unlimited", and require an explicit sentinel such as `MaxSize: -1` to
|
||
opt into unbounded. Validate `Options` in `NewMemoryProvider` and log the
|
||
effective values once at startup.
|
||
|
||
---
|
||
|
||
### 23. Low — constructors mutate the caller's config struct
|
||
|
||
`NewMemcacheProvider` writes `config.Servers`, `config.MaxIdleConns`,
|
||
`config.Timeout` and `config.Options` (`:41-57`); `NewRedisProvider` writes
|
||
`config.Host`, `config.Port`, `config.PoolSize`, `config.Options` (`:48-62`).
|
||
Both also assign into the `config == nil` replacement, which is at least local.
|
||
|
||
**Failure scenario.** A caller holds one config struct and constructs two
|
||
providers from it (a common test pattern, or a primary/replica setup). The second
|
||
construction sees the first one's defaults already applied, so "zero means
|
||
default" no longer holds and an intentional later change is silently ignored.
|
||
Callers reasonably assume a constructor does not write to their arguments.
|
||
|
||
**Recommendation.** Copy first: `cfg := *config` (plus a copy of `Options` if it
|
||
is non-nil, since it is a pointer), then apply defaults to the local copy.
|
||
|
||
---
|
||
|
||
### 24. Low — no defensive copy of `[]byte` in the memory provider
|
||
|
||
`Set` stores the caller's slice directly (`provider_memory.go:112`) and `Get`
|
||
returns the stored slice directly (`:77`, `:87`). Both share backing memory with
|
||
the caller.
|
||
|
||
**Failure scenario.** Two requests `GetBytes` the same key and receive the same
|
||
underlying array. If either mutates it in place — or if a caller reuses a buffer
|
||
it passed to `SetBytes` — the cached value changes for everyone, with no lock
|
||
held and no copy. For the session cache that means one request's scratch buffer
|
||
can rewrite another user's cached `UserContext`. Today every consumer goes
|
||
through `json.Marshal`/`json.Unmarshal` in `cache_manager.go`, which allocates
|
||
fresh slices, so this is latent rather than live; the `SetBytes`/`GetBytes`
|
||
API (`:56`, `:37`) exposes it directly to any future caller.
|
||
|
||
**Recommendation.** Copy on both sides in `MemoryProvider` — the Redis and
|
||
memcache providers get copies for free because the data crosses a socket, so
|
||
this also removes a behavioural difference between providers:
|
||
|
||
```go
|
||
buf := make([]byte, len(value))
|
||
copy(buf, value)
|
||
```
|
||
|
||
Document the ownership rule on the `Provider` interface either way.
|
||
|
||
---
|
||
|
||
### 25. Low — `example_usage.go` is a library file full of `log.Fatal`
|
||
|
||
`pkg/cache/example_usage.go` (266 lines) is compiled into the package and calls
|
||
`log.Fatal` at seventeen sites (`:18`, `:36`, `:44`, `:57`, `:64`, `:83`, `:96`,
|
||
`:103`, `:112`, `:127`, `:146`, `:154`, `:168`, `:187`, `:199`, `:224`, `:245`).
|
||
`log.Fatal` calls `os.Exit(1)`.
|
||
|
||
**Failure scenario.** These are exported functions (`ExampleInMemoryCache`,
|
||
`ExampleRedisCache`, `ExampleMemcacheCache`, …) with no `_test.go` suffix, so
|
||
they are part of the package's public API. Anything that calls one — a
|
||
misremembered name, a code-completion accident, a copied snippet — can terminate
|
||
the host process on a cache error, bypassing every panic handler and graceful
|
||
shutdown path. They also drag `log` into the package's dependency set while the
|
||
package deliberately imports no logger.
|
||
|
||
**Recommendation.** Move the file to `example_usage_test.go` (Go's testable-example
|
||
convention, which also makes them compile-checked and runnable), or to a
|
||
`_examples/` directory outside the package. Replace `log.Fatal` with returned
|
||
errors.
|
||
|
||
---
|
||
|
||
## What looks right
|
||
|
||
- `Provider` (`provider.go:9-44`) is a clean, minimal interface; the three
|
||
backends are genuinely swappable and the package has no import cycle problems
|
||
(it imports nothing from `ResolveSpec` at all).
|
||
- `hits`/`misses` use `atomic.Int64` (`provider_memory.go:35-36`) and are read
|
||
with `Load()` in `Stats`, so the counters themselves are race-free.
|
||
- `MemoryProvider.Delete` (`:173-191`) and `SetWithTags` (`:141-151`) do maintain
|
||
`tagToKeys` correctly, including deleting the tag entry when its key set
|
||
empties — the bug in finding 4 is the *other* paths not doing the same.
|
||
- `DeleteByTag` (`:194-228`) correctly handles multi-tag items: it strips only
|
||
the invalidated tag and keeps the item alive if other tags remain.
|
||
- `RedisProvider.DeleteByPattern` (`:196-225`) batches `DEL` in groups of 100
|
||
rather than buffering an unbounded pipeline, and checks `iter.Err()`.
|
||
- `RedisProvider.Close` (`:242-244`) correctly delegates to `client.Close()`.
|
||
- Both network providers verify connectivity at construction
|
||
(`provider_redis.go:75`, `provider_memcache.go:64`) and return an error rather
|
||
than deferring the failure to the first request — though the Redis `Ping` uses
|
||
a 5s blocking timeout on a `context.Background()`, which delays startup.
|
||
- Cache keys for query totals are SHA-256 hashes of the query shape
|
||
(`restheadspec/cache_helpers.go:88-92`), not raw user input — which is why the
|
||
key-injection risk in findings 9 and 12 is confined to the `pkg/security`
|
||
session path.
|
||
- `pkg/security/keystore_database.go:287` already keys on a hash
|
||
(`keystoreCacheKey(hash)`), which is the pattern `providers.go:398` should
|
||
follow.
|
||
|
||
## Suggested follow-up
|
||
|
||
Ordered by value:
|
||
|
||
1. **Make cache failures non-fatal in the auth path** (findings 1, 2, 9). This
|
||
is the single highest-value change: a cache problem should never be able to
|
||
reject a valid credential, and a revocation that cannot be honoured must be
|
||
loud.
|
||
2. **Key the session cache on `sha256(token)`** (findings 7, 9, 12) in
|
||
`pkg/security/providers.go:318`, `:398`, `:466`. Removes attacker control of
|
||
cache keys, bounds their length, and keeps the credential out of error
|
||
strings.
|
||
3. **Replace `DeleteByPattern`-based revocation with `DeleteByTag`**
|
||
(findings 2, 6), and make `ClearUserCache` actually scoped to the user.
|
||
4. **Fix the `defaultCache` race** (finding 3) with `atomic.Pointer` +
|
||
`sync.Once`, and close the displaced provider on swap (finding 14).
|
||
5. **Add `go test -race ./pkg/cache/...` to CI.** Findings 3 and 10 are
|
||
detectable in minutes with a small concurrent test; see
|
||
`_CROSS-CUTTING.audit.md`. `pkg/cache` currently has 69 lines of tests — two
|
||
functions, `TestSetDefaultCache` (`cache_test.go:9`) and
|
||
`TestGetDefaultCacheInitialization` (`:50`) — and is not in the package list
|
||
that `Makefile`/`.github/workflows/tests.yml` run.
|
||
6. **Centralize tag-index maintenance in `MemoryProvider`** (finding 4) and reset
|
||
it in `Clear`; add an eviction path that cannot leak.
|
||
7. **Make `MemoryProvider.Get` read-only** (finding 5) by moving `LastAccess`
|
||
and `HitCount` to atomics.
|
||
8. **Add single-flight to `GetOrSet`** (finding 8).
|
||
9. **Add TLS to `RedisConfig`/`MemcacheConfig`** (finding 15) and default it on.
|
||
10. **Give the package a logger and error metrics** (finding 17). Right now
|
||
`pkg/cache` cannot tell anyone that anything went wrong: 1 538 lines, zero
|
||
log statements, zero `recover()`, and eight `_ =` error discards outside the
|
||
examples file — seven of them in `provider_memcache.go`, precisely where the
|
||
tag index breaks (finding 12).
|
||
11. **Scope `Clear`** (finding 13) so it cannot flush the event broker's Redis DB.
|
||
|
||
## Cross-references
|
||
|
||
- `audit/pkg/logger.audit.md` — finding 2 (messages forwarded to Sentry
|
||
unscrubbed) is what makes finding 7 here dangerous; finding 4 (`CatchPanic`
|
||
swallows) is why findings 11/16/19 should not simply add `CatchPanic`.
|
||
- `audit/pkg/config.audit.md` — finding 3 (insecure transport defaults) is the
|
||
same theme as finding 15; `cache.redis.*` defaults live at
|
||
`pkg/config/manager.go:195-202` and expose no TLS field.
|
||
- `audit/pkg/security.audit.md` — findings 1, 2, 7, 8, 9 all land in
|
||
`pkg/security/providers.go`; the unbounded `go a.updateSessionActivity(...)`
|
||
per request (`:447`) is raised there.
|
||
- `audit/pkg/restheadspec.audit.md` — the query-total cache path
|
||
(`handler.go:829-860`) and tag-based invalidation (`:1477`, `:1711`, `:1785`,
|
||
`:1859`, `:1919`, `:2024`) are the other consumer; error returns from
|
||
`invalidateCacheForTags` are the invalidation-failure signal.
|
||
- `audit/pkg/metrics.audit.md` — `Provider` has `RecordCacheHit`/
|
||
`RecordCacheMiss`/`UpdateCacheSize` but no cache-error counter, and
|
||
`pkg/cache` calls none of them.
|
||
- `audit/pkg/_CROSS-CUTTING.audit.md` — no `-race` in CI; only
|
||
`pkg/resolvespec` and `pkg/restheadspec` are tested.
|