Willett AVTManufacturing Production and Technology Consulting
contact@willettavt.com360-303-2625
← Software & Systems

Grow the cluster, watch it migrate live.

Peanut is the resilience and scale layer for Walter: it routes uploads and reads to the right shard, watches every node's health, and runs rebalances without downtime.

A single Walter instance is already a complete file store, but it's also a single point of failure and a hard ceiling on capacity. Scaling that by hand, deciding which server holds which shard, noticing when one goes down, moving data around without breaking clients mid-migration, is exactly the kind of operational work nobody wants to do under pressure. Peanut is the piece that does it: a small, highly-available brain in front of the cluster that an operator, or an automated governance layer acting on their behalf, can drive through a handful of deliberate actions.

Peanut's topology screen listing every registered server with its shard, role, upstream replication partner, and up/down status, plus a form to register a new server

Every server, its shard, its role, and whether it's up, at a glance.

Growing the cluster is a real workflow, not a switch. Pick a new target shard count and start a rebalance, migrate the data over in controlled batches, then cut over once migration is complete. The riskier path, cutting over anyway with migration still incomplete, is there for the rare case that calls for it, but it sits behind its own explicit warning and a checkbox acknowledging the risk, and orphan cleanup afterward is its own deliberate step rather than something invisible.

Peanut's rebalancing screen with controls to start a rebalance, migrate in batches, cut over, force a cutover behind a warning and checkbox, and sweep orphans afterward

A rebalance you drive in steps: migrate, cut over, sweep, not one switch.

Key compliance is enforced, not just propagated. Publisher API keys are managed from one place, created, revoked, and re-pushed to every master at once, and a master that hasn't yet caught up on the latest key state is excluded from taking uploads until it does. A revoked key doesn't work its way out of the cluster eventually; there's no window where a stale node can still accept a write with a key that's already been pulled. A reconciliation sweep runs on demand too, copying any file missing from a server to the rest of its cluster.

Peanut's reconciliation screen with a sweep to copy missing files across the cluster, and a key-state table showing which masters are compliant

Compliance checked per master, not assumed once a key is pushed.

Underneath all of it is one rule: exactly one Peanut instance is ever active, the standby is a pure hot spare sharing a floating address with it. Every health check and routing decision has a single, unambiguous source of truth, so there's never a "whose opinion wins" problem to resolve. One failed reach is enough to mark a node down, one success is enough to mark it back up. No quorum, no debounce window to tune, and no distributed consensus protocol standing in for what a single decision-maker settles by construction.

What it does

Want a storage cluster that scales without anyone playing traffic cop by hand?