Valkey Instance Memory Limits: How to Right-Size a 256 MiB–2 GiB Cache Tier

Right-sizing a managed cache requires understanding that your raw dataset footprint does not equal your memory ceiling. When configuring Valkey instance memory limits between 256 MiB and 2 GiB, structural metadata, jemalloc allocation alignment, and client I/O buffers quickly consume a substantial portion of your configured maxmemory ceiling.

Selecting the right tier ensures your application does not experience unexpected out-of-memory errors or aggressive key evictions while avoiding unneeded infrastructure spend. This guide walks through the underlying memory accounting mechanics in Valkey, models eviction and connection overhead, and provides concrete sizing math for cache, session, and rate-limiting workloads.

Why Your Working Set Is Not Your Memory Limit

In practice, Valkey instance memory limits represent a hard ceiling for the entire process keyspace, dictionary metadata, and connection states—not a clean user-data quota.

Every stored key incurs internal bookkeeping costs. When you execute SET user:10492 session_blob, the engine does not merely write the bytes of the key name and the serialized payload to a contiguous byte array. Instead, Valkey allocates:

  • A dictionary entry (dictEntry) inside the internal hash table pointing to the key and value.
  • A robj (Valkey object metadata structure) tracking object type, encoding, and LRU/LFU bits.
  • An SDS (Simple Dynamic String) header for both the key and the string payload.
  • Memory allocator padding: because underlying allocators such as jemalloc group allocations into fixed power-of-two size classes, a 33-byte string often claims a 48- or 64-byte bucket.

Beyond key-level storage, memory budgets must absorb Valkey memory overhead from internal replication buffers, client input/output buffers, and hash table rehashing churn. A Valkey 256 MiB limit provides 268,435,456 bytes total. If your application pushes 200 MiB of serialized payload data across hundreds of thousands of individual keys, the engine will cross that ceiling rapidly.

Because internal metadata and allocator padding depend heavily on key cardinality and payload length, planning instance headroom requires accounting for more than raw payload volume. Consider a SaaS session cache: storing 200,000 active sessions with an average payload size of 400 bytes alongside a 60-byte key name totals roughly 92 MB of raw application strings. Once you account for dictionary overhead, SDS structures, object headers, and jemalloc bin rounding, the real memory consumption frequently lands between 140 MB and 165 MB—before accounting for client connection buffers.

How Valkey Counts Memory: used_memory, maxmemory and the Overhead Line Items

Evaluating memory consumption through the Redis-compatible INFO memory command reveals multiple distinct metrics. Misinterpreting these fields leads to incorrect sizing assumptions:

  • used_memory: Total bytes allocated by Valkey using its internal allocator (jemalloc). This includes data, internal dictionaries, and client output buffers. This metric directly governs when eviction triggers against the configured maxmemory ceiling.
  • used_memory_rss: Resident Set Size—the actual physical RAM allocated to the Valkey process by the operating system.
  • used_memory_dataset: The memory consumed by key names, values, and secondary data structures, excluding engine overhead.
  • used_memory_overhead: Memory consumed by basic engine internals, client connections, replication backlogs, and global dictionary tables.
  • mem_fragmentation_ratio: The ratio of used_memory_rss to used_memory. A ratio significantly above 1.5 indicates that memory fragmentation is stranding memory at the OS level that the allocator cannot reuse immediately.

When evaluating providers, the billed metric or enforced limit is generally tied to maxmemory. However, if used_memory_rss climbs sharply above maxmemory due to memory fragmentation or connection churn, an instance risks kernel-level out-of-memory (OOM) termination if total system limits are breached.

Per-key overhead also changes based on data structures. Small string keys incur high percentage overhead due to fixed dictionary headers. Conversely, hashes, lists, and sorted sets containing small elements benefit from compact memory encodings (such as listpacks). For example, storing 100 fields inside a single hash yields lower overhead per field than storing 100 separate string keys, because the compact encoding avoids individual dictEntry allocations.

During the first week after a migration or launch, inspect these metrics daily. If used_memory routinely approaches your tier's maxmemory ceiling, or if used_memory_overhead accounts for a substantial portion of overall consumption, evaluate key length normalization or plan an upgrade to the next tier.

Eviction Behavior at the Limit: What Happens When You Hit the Ceiling

When used_memory reaches your configured ceiling, the instance executes its configured Valkey key eviction documentation policy. Valkey eviction behavior varies substantially across different policies:

  • noeviction: The engine refuses to evict keys. Read commands and deletes continue to execute, but commands requesting memory allocation (such as SET, HSET, or LPUSH) immediately return OOM command not allowed when used memory > 'maxmemory' errors back to your application.
  • allkeys-lru: The engine samples keys across the entire database and evicts the least used keys, regardless of whether an explicit TTL was defined.
  • volatile-lru: The engine evicts least used keys only from the subset of keys that have an explicit expiration set. Keys lacking an expiry are retained.
  • allkeys-lfu / volatile-lfu: The engine evicts keys based on least frequent access frequency rather than elapsed idle time.

The operational failure modes between these policies are distinct. An unhandled noeviction state surfaces as hard exceptions in your application logs, breaking API requests and background workers. Conversely, aggressive eviction under allkeys-lru manifests as silent cache misses. Your origin database suddenly absorbs full read traffic, resulting in increased database query latency while the cache appears online.

Using volatile-lru on a store containing mixed key types introduces subtle failure modes. If your session keys include an expiration, but feature flags or configuration keys do not, memory pressure forces the engine to evict session keys prematurely. Once all volatile keys are removed, any further writes to non-volatile keys fail with OOM errors. For a dedicated cache tier, allkeys-lru provides predictable behavior. For rate limiters and session stores, strict TTL enforcement is mandatory so expired keys release allocations before evictions disrupt active users.

Sizing Three Real Workloads: Cache, Sessions and Rate Limiting

Right-sizing depends entirely on access patterns and data lifecycles. Here is how to model three standard workloads against tiers spanning 256 MiB to 2 GiB.

1. Read Cache: Size to the Working Set, Not Origin Tables

You do not need to cache your entire database. If an origin table occupies 8 GiB on disk, but the vast majority of active queries target only recent records, the hot working set may represent only a small fraction of the table—such as 350 MB of serialized JSON payloads.

In an application where dictionary structures, allocator bin rounding, and client buffers require substantial overhead beyond raw strings, a scenario might model:

350 MB raw data + estimated metadata and buffer overhead = ~525 MB allocated memory

A 1 GiB instance provides 1,024 MiB of ceiling, granting roughly 499 MiB of breathing room for query spikes, connection buffers, and key cardinality growth.

2. Session Store: Exact Record Sizing and Restart Risk

Session stores have measurable record shapes. Sizing can be calculated using active user counts and concurrency factors:

  • Active authenticated sessions: 150,000
  • Key name format: sess:usr:<uuid> (~45 bytes)
  • Serialized session payload: 600 bytes
  • Total raw payload: 150,000 × (45 + 600) bytes = 96.75 MB
  • Jemalloc allocation rounding and dictionary overhead: ~154.8 MB
  • Headroom buffer (allowing breathing room for concurrent write bursts): ~201.2 MB

This workload fits inside a Valkey 256 MiB limit. However, review your application's restart resiliency. Cache data can be lost on restart; sessions may require reauthentication if stored without durable secondary mechanisms. Steada is for cache, sessions, rate limiting, and low-risk metadata that can roll back — not source-of-truth data without an independent recovery path. If your session store drops during a maintenance restart, your frontend must handle re-authentication cleanly without overwhelming auth endpoints.

3. API Rate Limiting: High Cardinality, Small Footprint

Rate limiting workloads typically track simple integers or sliding window sorted sets. For fixed-window counters:

  • Unique active API clients per minute: 400,000
  • Key structure: rate:<client_id>:<minute_timestamp> (~35 bytes)
  • Value: integer counter (encoded as an int inside the robj, claiming minimal byte storage)
  • Raw string data: ~14 MB
  • Internal dictionary entries: 400,000 keys × ~56 bytes per dictEntry = ~22.4 MB
  • Jemalloc metadata overhead: ~20 MB
  • Total used memory: ~56.4 MB

While 56.4 MB easily fits within 256 MiB, the challenge is key churn. If TTLs are omitted or prolonged, 400,000 keys per minute will consume roughly 56 MB every 60 seconds, exhausting 1 GiB of memory in under twenty minutes. Strict TTL hygiene is required when running rate limiters at scale.

Connection Ceilings and Memory: The Limit You Hit First

Memory allocations are not solely dictated by stored keys. Valkey allocates memory buffers for every connected client. In microservice environments or serverless deployments, connection-related memory consumption can outpace key data.

Every active connection maintains an input buffer (which dynamically scales up to several megabytes to parse incoming payloads) and an output buffer (which queues responses if the client reads slowly). In standard operation, an idle connection claims roughly 20 KiB to 40 KiB of process memory. However, under high concurrency, active pipelines, or slow networks, client output buffers can expand significantly.

500 idle connections × 30 KiB = 15 MB
500 active connections under burst (e.g., 256 KiB average buffer) = 128 MB

On a 256 MiB instance, 128 MB of transient buffer allocations leaves only 128 MiB for all stored data. If your keyspace already occupies 180 MiB, connection growth will immediately push the instance past its maxmemory ceiling, triggering unwanted key evictions or client disconnections.

Application architecture determines connection footprints:

  • Application Connection Pools: 15 web application nodes, each maintaining a connection pool of 20 connections, creates 300 steady-state connections. This creates a predictable memory baseline.
  • Serverless Functions (FaaS): 600 concurrent lambda executions each creating a fresh connection on invocation can produce 600 simultaneous connections, exhausting both memory buffers and engine concurrency limits.

To prevent buffer-driven memory exhaustion:

  1. Set conservative client pool maximums in your web frameworks (for instance, 10–25 connections per process instead of 100).
  2. Ensure client output buffer limits (client-output-buffer-limit) close delinquent clients that fail to consume queued responses.
  3. Implement reconnect backoffs with jitter to prevent connection storms when an instance restarts.

Choosing a Tier: 256 MiB, 512 MiB, 1 GiB or 2 GiB

When selecting a managed tier, evaluate both memory requirements and baseline workload economics. Steada charges a flat monthly price per plan; cost does not scale per request or per command, which is the explicit contrast with request-metered providers.

Steada is a cost-first managed Valkey service — a Redis-compatible, BSD-licensed in-memory key-value store — for cost-sensitive production teams. Steada is independent of the Valkey project and the Linux Foundation. The tenant data plane runs in DigitalOcean NYC3, with one Valkey instance per database, TLS endpoints, scoped credentials and memory limits. Steada does not offer multi-region or active-active replication, and Steada does not offer a formal SLA or uptime guarantee.

As published on the Steada pricing page, self-service monthly plan tiers include:

  • Starter (256 MiB) at $49/month: Realistic usable payload space is ~140–170 MB. Ideal for focused rate limiting, small authentication stores, or microservice edge caches.
  • Growth (512 MiB) at $89/month: Realistic usable payload space is ~300–360 MB. Accommodates medium session workloads or hot query caches for small SaaS platforms.
  • Scale (1 GiB) at $149/month: Realistic usable payload space is ~650–750 MB. Suitable for multi-table database caching with active working sets under high command volume.
  • Scale+ (2 GiB) at $249/month: Realistic usable payload space is ~1.3–1.5 GB. Handles large session pools and memory-intensive caching topologies.

Check the Steada pricing page before provisioning. Nano and Micro configurations are not offered via self-service Checkout.

The table below summarizes usable capacities across these four tiers:

Steada Tier Configured maxmemory Estimated Usable Payload Capacity Recommended Workload Profile
Starter ($49/mo) 256 MiB ~140–170 MB API rate limiters, transient microservice caches, lightweight sessions (<100k users).
Growth ($89/mo) 512 MiB ~300–360 MB Active web session stores, targeted relational query caching, worker status stores.
Scale ($149/mo) 1 GiB ~650–750 MB High-throughput hot set caching, medium-sized session catalogs, high-concurrency apps.
Scale+ ($249/mo) 2 GiB ~1.3–1.5 GB Extensive query caching, large user-state graphs, high connection concurrency pools.

Workload volume alters cost tradeoffs. Metered providers such as Upstash offer pay-as-you-go models alongside plan options with explicit capacity boundaries (see Upstash pricing). Low-volume workloads executing occasional lookups can cost less on PAYG. However, SaaS systems running sustained, command-heavy workloads accumulate substantial variable usage charges on per-request models. You can evaluate your access profile using the pricing calculator.

If your dataset outgrows its allocated boundaries, an in-place resize can interrupt in-flight network connections. Sizing with moderate memory headroom avoids frequent operational resizing cycles.

Measuring Before You Commit: Telemetry, Alerts and Export

Do not guess your memory overhead—measure real workloads in staging or production. Steada includes per-database usage telemetry, percentile latency, a projected month-end cost labeled a hypothesis, native threshold alerting, and read-only Prometheus + CSV export on the same tier.

Run a two-week capacity review tracking four primary metrics:

  1. Peak used_memory vs. maxmemory: Observe your utilization curve during peak traffic hours. If utilization consistently approaches your ceiling, plan for the next tier before writes fail or evictions discard active working sets.
  2. Evicted Keys Count: Query INFO stats to monitor evicted_keys. For a dedicated cache running an LRU policy, occasional evictions indicate normal operation. For rate limiters or session stores, any non-zero rate of eviction signifies premature data expulsion.
  3. Connection High-Water Mark: Track connected_clients and blocked_clients alongside used_memory_overhead. Spikes in connected clients that correspond to memory increases point to oversized client connection pools.
  4. Percentile Latency (p99): Measure command response latency. Sharp increases in p99 latency often occur when Valkey spends CPU cycles traversing large hash tables to process evictions under severe memory constraints.

Note that usage telemetry export records instance operational metrics, not database key contents. Steada has no completed compliance certifications (SOC 2, HIPAA, PCI, ISO 27001) today, and Steada makes no regulated-data commitments; do not store regulated or protected data such as PHI.

What Changes When You Migrate: Client Behavior and Rollback

Valkey is a Redis-compatible open-source fork using the standard RESP (REdis Serialization Protocol) wire format. The default connection path is native Redis/Valkey RESP over TLS with password authentication. Moving existing Node.js (ioredis), Python (redis-py), or Go (go-redis) clients requires updating your connection string and enabling TLS.

When migrating clients, observe two protocol-level details:

  • Supported Command Set: While Valkey supports standard Redis key-value, hash, set, and transaction commands, compatibility is not universal. Steada does not support Redis modules such as RediSearch, RedisJSON, or RedisBloom. Validate your application's command requirements against our documented compatibility subset prior to cutover.
  • Connection Protocols: Steada does not claim full Upstash REST API parity; the default path is native RESP over TLS, with only a narrow REST compatibility preview. If your architecture relies on HTTP REST requests to execute commands, verify your client library path or update your code to standard RESP over TLS via our connection documentation.

Plan your operational failover and rollback architecture around the volatility of in-memory caching. Single-instance managed caches running without automatic replica failover can experience restarts during host maintenance or configuration resizing. The following diagram illustrates the degraded fallback paths your application should execute during connection drops or restart recovery:

Application Client Layer RESP / TLS Valkey Cache 256 MiB – 2 GiB Eviction / Restart Cache Hit Fallback: DB Read + Circuit Breaker Origin Database Relational / Persistent Absorbs misses Rate-limits concurrency

Before initiating DNS cutover in production, run an automated restart test in staging. Verify that your application's connection pool retries with exponential backoff rather than spamming reconnect requests, that database query throttling protects backend stores during a cold cache warm-up, and that session disconnects prompt clean re-authentication flows.

Common Sizing Mistakes and How to Avoid Them

Engineers sizing a cache tier for the first time often encounter predictable memory management pitfalls:

  • Sizing for the Full Database: Caches should hold active working sets, not mirror disk storage. If you load multi-gigabyte historical records into a 512 MiB instance, Valkey will immediately evict active data to retain cold records. Restrict your caching layer to frequently queried items.
  • Neglecting Key Name Length: At high cardinality, key names consume significant allocations. Storing 3,000,000 keys using the key pattern organization:enterprise:customer:session:token:<uuid> uses roughly 180 MB of memory just for key identifiers. Shortening prefixes to org:sess:<uuid> saves tens of megabytes of process memory.
  • Omitting TTLs on Ephemeral Keys: Creating keys via SET or HSET without setting an explicit expiration will leave them in memory indefinitely. Under a volatile-lru eviction policy, untagged keys cannot be evicted, eventually triggering out-of-memory errors on subsequent writes.
  • Sizing should account for expected growth rather than operating near full capacity week-to-week.
  • Diagnosing Latency by Upgrading Memory: When command latency spikes, engineers often assume memory exhaustion is the root cause. Spikes in p99 latency are frequently caused by blocking commands (such as KEYS * or large SMEMBERS operations) or unoptimized client pooling rather than insufficient maxmemory. Inspect command distributions before purchasing a larger tier.

Frequently Asked Questions About Valkey Instance Memory Limits

How much of a 256 MiB Valkey instance is actually available for my keys?

Between 140 MiB and 170 MiB. A Valkey 256 MiB limit provides 268,435,456 bytes for total process memory. Internal dictionary hash tables, robj wrappers, jemalloc allocator alignment padding, and client connection buffers consume the remaining capacity. Sizing must account for data structure headers, allocator bin rounding, and connection buffers in addition to raw payload bytes.

What happens when a Valkey instance hits its memory limit?

Behavior depends on the configured maxmemory-policy. Under noeviction, Valkey rejects new write commands with an OOM error while continuing to serve read and delete requests. Under allkeys-lru or allkeys-lfu, the instance silently evicts keys to allocate memory for incoming writes. Under volatile-* policies, it evicts only keys with an explicit expiration.

Which maxmemory-policy should I use for a cache versus a session store?

For a pure rebuildable cache, use allkeys-lru. This ensures the least accessed items drop out cleanly to make room for active data without rejecting writes. For a session store or rate limiter where items carry explicit expirations, use volatile-lru or noeviction with strict TTL enforcement to prevent unexpected session terminations.

Do connections count against the memory limit?

Yes. Every connected client claims memory for socket input buffers and output reply buffers, typically starting at 20 KiB to 40 KiB when idle and expanding to hundreds of kilobytes during active query bursts. Having 500 active connections can easily consume 50 MiB to 100 MiB of your memory limit regardless of keyspace size.

Can I resize a Valkey database without downtime?

No. Applications must support transient reconnects using exponential backoff and handle cold cache warm-up states cleanly when resizing a tier.

Once you have a measured peak used_memory number, run it through the pricing calculator or view plan tiers on the pricing page to see which tier fits. You can also explore the live dashboard preview or verify your command set against our compatibility documentation before provisioning at https://steada.dev/start/.