SahyadriX
Sahyadri's memory-hard Proof-of-Work — cSHAKE256 bound to a 16 MB random-access loop, engineered to compress the gap between ASICs and commodity hardware.
Overview
SahyadriX is the application-layer Proof-of-Work function used by the Sahyadri network. It combines cSHAKE256 (NIST FIPS 202) with an 8-stage XOR memory loop operating over a 16 MB random-access memory pad. Every hash requires the full 16 MB working set to be resident in memory — making ASIC implementations economically unattractive and preserving broad participation across CPUs and GPUs.
Unlike PoW algorithms optimized for pure arithmetic throughput (SHA-256, kHeavyHash), SahyadriX is memory-latency bound. Speed depends on how fast a device can perform random reads from a 16 MB buffer — not on how many hash units fit on a die.
Algorithm
Each evaluation proceeds in three stages. The middle stage is the ASIC-killer.
Stage 1 — Initial Hash
The input (pre-PoW block hash, timestamp, nonce) is hashed with cSHAKE256, producing a 32-byte seed.
Stage 2 — Memory Pad Fill
The 16 MB buffer is deterministically filled from the seed. Each 32-byte window is produced by hashing the previous window together with its offset:
seed_hash = cSHAKE256(input)
for i in (0 .. MEM_SIZE).step_by(WINDOW) {
seed_hash = cSHAKE256(seed_hash || i.to_le_bytes())
mem_pad[i..i+WINDOW].copy_from_slice(seed_hash)
}Stage 3 — Random-Access Mixing
1024 rounds of pseudo-random XOR reads from the memory pad. Two offsets are derived from the current hash and used to mix two 32-byte windows into the next hash:
for _ in 0..ROUNDS {
idx1 = u32::from_le_bytes(current_hash[0..4]) % MAX_INDEX
idx2 = u32::from_le_bytes(current_hash[4..8]) % MAX_INDEX
current_hash = cSHAKE256(
current_hash || mem_pad[idx1..idx1+WINDOW] || mem_pad[idx2..idx2+WINDOW]
)
}The offsets are data-dependent: the location of the next read is unpredictable until the previous hash completes. An ASIC cannot pipeline these reads — it must serialize on memory latency, exactly like a CPU. This is what collapses the ASIC advantage.
Constants
| Parameter | Value | Purpose |
|---|---|---|
MEM_SIZE | 16 MB | Working set required per evaluation |
ROUNDS | 1024 | Number of random-access mixing rounds |
WINDOW | 32 bytes | Bytes read per memory access |
| Hash primitive | cSHAKE256 | NIST FIPS 202, quantum-resistant |
Memory-Hard Design
The 16 MB constant was chosen to sit at the intersection of three constraints: CPU L3 cache sizes, ASIC manufacturing costs, and consumer-grade hardware accessibility.
L3 Cache Fit
Modern CPUs have L3 caches between 4 MB (low-end) and 96 MB (high-end). A 16 MB working set:
- Fits mid-to-high-end CPUs (Intel i5 12th gen+, Ryzen 5+) — no RAM round-trips needed
- Spills partially on low-end CPUs (Celeron, older i3) — still competitive, just slower
- Dominates GPU workloads — GDDR6 bandwidth matters more than shader count
A 32 MB working set would exclude all CPUs with less than 32 MB L3. That is a meaningful fraction of the consumer market. A 16 MB set keeps them in.
ASIC Cost Floor
The 16 MB requirement imposes a hard cost floor on any specialized hardware. Fast on-die SRAM costs approximately $50–100 per megabyte in volume. That puts the memory alone at $800–1600 per chip — before adding compute, I/O, power delivery, and packaging.
| Memory Required | SRAM Cost Floor | Realistic ASIC Cost |
|---|---|---|
| 16 MB | ~$800–1600 | $2,000+ |
| 32 MB | ~$1600–3200 | $4,000+ |
| 64 MB | ~$3200–6400 | $8,000+ |
| 256 MB (RandomX-class) | ~$12,800–25,600 | $25,000+ |
The result is a nontrivial economic barrier — high enough to deter small attackers, low enough that CPUs and GPUs are not pushed out.
Device Balance
SahyadriX does not aim to make ASICs impossible. It aims to keep them close enough to commodity hardware that a mining advantage does not become a mining monopoly.
| Device | Memory | Relative Throughput | Notes |
|---|---|---|---|
| ASIC (theoretical) | On-chip, expensive | ~2–3× GPU | Constrained by 16 MB requirement; cost floor $2K+ |
| GPU (consumer) | 4–24 GB VRAM | Baseline | High parallelism, but memory-latency bound |
| CPU (modern, mid-range) | 8+ MB L3 | ~0.3–0.5× GPU | Single-thread, memory-latency optimised |
| CPU (high-end) | 32+ MB L3 | ~0.8–1.2× GPU | Entire memory pad fits in L3 |
A mid-range laptop CPU stays within 2–3× of a gaming GPU. That margin is small enough that mining remains a hobbyist activity rather than an industrial one.
Post-Quantum Layer
SahyadriX is bound into a fully post-quantum stack:
| Component | Algorithm | Quantum Security |
|---|---|---|
| PoW evaluation | cSHAKE256 | 256-bit preimage |
| Block hashing | SHA3-256 | 256-bit preimage |
| Merkle commitment | SHA3-256 | 128-bit collision (85-bit under BHT) |
| Signatures | ML-DSA-65 | NIST Level 3 |
No elliptic-curve cryptography appears anywhere in the PoW or block verification path. A quantum adversary gains no meaningful advantage.
PoW Bound to Block Contents
SahyadriX is cryptographically bound to the exact block template. Any change to block contents or the header's committed state root invalidates the nonce and forces a full recomputation. This prevents precomputation attacks where a miner might try to reuse work across template variations.
pub fn finalize_with_nonce(&self, nonce: u64) -> Hash {
let mut data = Vec::with_capacity(48);
data.extend_from_slice(self.pre_pow_hash.as_bytes()); // block contents
data.extend_from_slice(&self.timestamp.to_le_bytes());
data.extend_from_slice(&nonce.to_le_bytes());
SahyadriX::hash(&data) // 16 MB memory-hard evaluation
}Verification
Verification re-evaluates the full SahyadriX hash — there is no shortcut, no fast-path, no low-memory verification. Every node and every light client performing full header validation runs the same 16 MB loop.
On a modern CPU, single-thread verification completes in approximately 5–15 ms. This is why header sync on resource-constrained devices remains practical: verification is bounded and predictable, unlike mining which is probabilistic.
Evolution Path
SahyadriX is designed to be upgradable. The memory size, round count, and hash window are all consensus parameters that can be raised through a coordinated hard fork. The whitepaper establishes three tiers:
| Version | Memory | Trigger Condition |
|---|---|---|
| v1 (current) | 16 MB | Genesis — broad CPU/GPU participation |
| v2 | 32 MB | GPU dominance emerges; mid-range CPUs becoming uncompetitive |
| v3 | 64 MB | Dedicated hardware (FPGA/ASIC) becomes economically visible |
| v4+ | 128 MB+ | Ongoing hardening against industrial-scale centralization |
Upgrades are reactive, not scheduled. The protocol changes only when the hardware landscape requires it. There is no fixed roadmap that forces migration; the current parameters remain stable until an upgrade becomes necessary.
Comparison with Other PoW Designs
| Algorithm | Memory | Primary Target | ASIC Resistance |
|---|---|---|---|
| SHA-256 (Bitcoin) | Small, on-chip | Raw compute | None — ASIC dominant |
| Ethash (Ethereum PoW) | ~1 GB DAG | GPU memory bandwidth | Partial — GPU-heavy |
| RandomX (Monero) | ~256 MB | CPU-centric | Strong — CPU-favored |
| kHeavyHash | ~32 MB | GPU arithmetic | Partial — GPU-favored |
| SahyadriX | 16 MB | Balanced CPU / GPU | Latency-bound, cost-floored |
SahyadriX is deliberately not RandomX (which excludes GPUs) and not kHeavyHash (which excludes CPUs). It sits in the middle: a working set small enough for CPU L3, large enough to matter for ASICs, and balanced enough that neither device class dominates.
Security Analysis
| Threat | Mitigation |
|---|---|
| ASIC centralization | 16 MB working set imposes $2K+ per-chip cost floor; ASIC advantage vs. GPU remains ~2–3×, not 100× |
| GPU centralization | Memory-latency bound mixing prevents pure arithmetic throughput from dominating; CPU L3-resident evaluation is competitive |
| Precomputation across templates | PoW is bound to block contents including the state root; any change forces full recomputation |
| Quantum adversary (Grover) | cSHAKE256 retains 128-bit preimage resistance under Grover; no ECC anywhere in the PoW path |
| Low-memory verification attacks | No shortcut — full re-evaluation required, no time-memory trade-off |
| Parameter ossification | Memory size, round count, and window are consensus parameters; upgradable via hard fork |
Summary
SahyadriX is a memory-hard, latency-bound Proof-of-Work designed for broad accessibility rather than absolute ASIC resistance. Its 16 MB working set fits modern CPU L3 caches, imposes a real cost floor on specialized hardware, and keeps GPU and CPU miners within a factor of a few. The algorithm is fully post-quantum, bound to block contents, and structured for future memory-size upgrades as the hardware landscape shifts.
The design accepts a simple truth: no PoW is permanently ASIC-proof. The goal is not to win that race, but to slow it down enough that mining remains a distributed activity rather than a specialized one. That is what SahyadriX optimises for — and what the 16 MB constant represents.