Implement Argon2id exactly as specified in RFC 9106 in pure Go with no external crypto libraries: BLAKE2b, the variable-length hash H', the 1024-byte compression G with BlaMka multiply-add, the p-lane memory fill (data-independent addressing for the first two slices of pass 0, Argon2d-style afterwards), and the final XOR + tag hash.
I've implemented and verified the solution. The complete write-up is saved at ~/argon2id/SOLUTION.md, and the runnable package lives in ~/argon2id/. Below is the solution.
Implement Argon2id exactly as specified in RFC 9106 in pure Go with no external crypto libraries: BLAKE2b, the variable-length hash H', the 1024-byte compression G with BlaMka multiply-add, the p-lane memory fill (data-independent addressing for the first two slices of pass 0, Argon2d-style afterwards), and the final XOR + tag hash.
Acceptance criteria: byte-exact RFC vectors (including t>1, p>1, secret key, associated data), peak allocation of 64 MiB at m=65536, correct 8 KiB/t=1 boundary, and rejection of out-of-range parameters.
G: the second message word goes into the second a, not into ca = a + b + x
d = rotr(d ^ a, 32)
c = c + d
b = rotr(b ^ c, 24)
a = a + b + y <-- y is added here
d = rotr(d ^ a, 16)
c = c + d
b = rotr(b ^ c, 63)
The classic mistranscription puts y into c = c + d. The trap: BLAKE2b("") still passes because an all-zero message makes the permutation message-independent, so BLAKE2b("abc") is the first vector that catches it. This was the first real bug found.
indexAlpha: J1² needs a 64-bit productx = (J1 * J1) >> 32
y = (reference_area_size * x) >> 32
z = reference_area_size - 1 - y
If J1 is a uint32, J1*J1 wraps modulo 2³² and >>32 is always 0, collapsing the reference-index distribution. Fix: x := (uint64(rand) * uint64(rand)) >> 32.
H0 binds requested m; the fill uses m' = 4·p·floor(m/(4p))Address blocks embed the rounded m' (the reference's memory_blocks). Mixing these breaks any m not a multiple of 4p.
Data-independent addressing applies only to pass == 0 && slice < 2, from counter blocks G(G(input,0)) where input = LE64(pass)||LE64(lane)||LE64(slice)||LE64(m')||LE64(t)||LE64(type)||LE64(counter). All other segments use the previous block's first word (J1|J2). In pass 0 / slice 0 the lane is forced to the current lane.
B[i][0] = H'^1024(H0||LE32(0)||LE32(i)), B[i][1] = H'^1024(H0||LE32(1)||LE32(i)); pass 0 / slice 0 starts at index 2.
H' for T > 64r = ceil(T/32)-2, output is V1[0:32] || … || Vr[0:32] || H^(T-32r)(Vr).
One []Block (Block = [128]uint64 = 1024 B) is 64 MiB at m=65536; H0/H' scratch is on the stack, so only the matrix and returned tag hit the heap.
go.mod:
module argon2id
go 1.26
blake2b.gopackage argon2id
import "encoding/binary"
var blake2bIV = [8]uint64{
0x6a09e667f3bcc908, 0xbb67ae8584caa73b,
0x3c6ef372fe94f82b, 0xa54ff53a5f1d36f1,
0x510e527fade682d1, 0x9b05688c2b3e6c1f,
0x1f83d9abfb41bd6b, 0x5be0cd19137e2179,
}
var blake2bSigma = [12][16]uint8{
{0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15},
{14, 10, 4, 8, 9, 15, 13, 6, 1, 12, 0, 2, 11, 7, 5, 3},
{11, 8, 12, 0, 5, 2, 15, 13, 10, 14, 3, 6, 7, 1, 9, 4},
{7, 9, 3, 1, 13, 12, 11, 14, 2, 6, 5, 10, 4, 0, 15, 8},
{9, 0, 5, 7, 2, 4, 10, 15, 14, 1, 11, 12, 6, 8, 3, 13},
{2, 12, 6, 10, 0, 11, 8, 3, 4, 13, 7, 5, 15, 14, 1, 9},
{12, 5, 1, 15, 14, 13, 4, 10, 0, 7, 6, 3, 9, 2, 8, 11},
{13, 11, 7, 14, 12, 1, 3, 9, 5, 0, 15, 4, 8, 6, 2, 10},
{6, 15, 14, 9, 11, 3, 0, 8, 12, 2, 13, 7, 1, 4, 10, 5},
{10, 2, 8, 4, 7, 6, 1, 5, 15, 11, 9, 14, 3, 12, 13, 0},
{0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15},
{14, 10, 4, 8, 9, 15, 13, 6, 1, 12, 0, 2, 11, 7, 5, 3},
}
func rotr64(x uint64, n uint) uint64 { return x>>n | x<<(64-n) }
func blake2bCompress(h *[8]uint64, block []byte, t uint64, final bool) {
var m [16]uint64
for i := range m {
m[i] = binary.LittleEndian.Uint64(block[i*8:])
}
var v [16]uint64
copy(v[:8], h[:])
copy(v[8:], blake2bIV[:])
v[12] ^= t
if final {
v[14] = ^v[14]
}
for r := 0; r < 12; r++ {
s := &blake2bSigma[r]
blake2bG(&v, 0, 4, 8, 12, m[s[0]], m[s[1]])
blake2bG(&v, 1, 5, 9, 13, m[s[2]], m[s[3]])
blake2bG(&v, 2, 6, 10, 14, m[s[4]], m[s[5]])
blake2bG(&v, 3, 7, 11, 15, m[s[6]], m[s[7]])
blake2bG(&v, 0, 5, 10, 15, m[s[8]], m[s[9]])
blake2bG(&v, 1, 6, 11, 12, m[s[10]], m[s[11]])
blake2bG(&v, 2, 7, 8, 13, m[s[12]], m[s[13]])
blake2bG(&v, 3, 4, 9, 14, m[s[14]], m[s[15]])
}
for i := 0; i < 8; i++ {
h[i] ^= v[i] ^ v[i+8]
}
}
func blake2bG(v *[16]uint64, a, b, c, d int, x, y uint64) {
v[a] += v[b] + x
v[d] = rotr64(v[d]^v[a], 32)
v[c] += v[d]
v[b] = rotr64(v[b]^v[c], 24)
v[a] += v[b] + y
v[d] = rotr64(v[d]^v[a], 16)
v[c] += v[d]
v[b] = rotr64(v[b]^v[c], 63)
}
func blake2b(out, in []byte) {
if len(out) == 0 || len(out) > 64 {
panic("argon2id: blake2b digest size out of range")
}
h := blake2bIV
h[0] ^= 0x01010000 ^ uint64(len(out))
var t uint64
for len(in) > 128 {
t += 128
blake2bCompress(&h, in[:128], t, false)
in = in[128:]
}
var last [128]byte
copy(last[:], in)
t += uint64(len(in))
blake2bCompress(&h, last[:], t, true)
for i := 0; i < len(out); i++ {
out[i] = byte(h[i>>3] >> (8 * uint(i&7)))
}
}
func blake2b512(in []byte) [64]byte {
var d [64]byte
blake2b(d[:], in)
return d
}
argon2id.go// Package argon2id implements Argon2id per RFC 9106 with only the Go standard library.
package argon2id
import (
"encoding/binary"
"errors"
"math"
)
const Version = 0x13
const (
blockWords = 128
blockBytes = 1024
syncPoints = 4
argon2idVal = 2
)
type Block [blockWords]uint64
type Params struct {
Parallelism uint32 // p, 1 <= p <= 2^24-1
Memory uint32 // m, KiB, >= 8*p
Iterations uint32 // t, >= 1
TagSize uint32 // T, >= 4
}
var (
ErrTimeTooSmall = errors.New("argon2id: iterations t must be >= 1")
ErrMemoryTooSmall = errors.New("argon2id: memory m must be >= 8*parallelism")
ErrLanesTooFew = errors.New("argon2id: parallelism p must be >= 1")
ErrLanesTooMany = errors.New("argon2id: parallelism p must be <= 2^24-1")
ErrTagTooShort = errors.New("argon2id: tag length must be >= 4")
ErrSaltTooShort = errors.New("argon2id: salt must be >= 8 bytes")
)
func DeriveKey(password, salt, secret, associatedData []byte, params Params) ([]byte, error) {
if err := params.validate(); err != nil {
return nil, err
}
if len(salt) < 8 {
return nil, ErrSaltTooShort
}
if uint64(len(password)) > math.MaxUint32 || uint64(len(salt)) > math.MaxUint32 ||
uint64(len(secret)) > math.MaxUint32 || uint64(len(associatedData)) > math.MaxUint32 {
return nil, errors.New("argon2id: input too long")
}
p, m, t, tagSize := params.Parallelism, params.Memory, params.Iterations, params.TagSize
h0 := initialHash(password, salt, secret, associatedData, p, m, t, tagSize)
memBlocks := (m / (syncPoints * p)) * (syncPoints * p) // m' >= 8p
laneLength := memBlocks / p
segmentLength := laneLength / syncPoints
mem := make([]Block, memBlocks)
initBlocks(mem, &h0, p, laneLength)
for pass := uint32(0); pass < t; pass++ {
for slice := uint32(0); slice < syncPoints; slice++ {
for lane := uint32(0); lane < p; lane++ {
fillSegment(mem, pass, lane, slice, p, laneLength, segmentLength, t, memBlocks)
}
}
}
return extractTag(mem, p, laneLength, tagSize), nil
}
func Argon2id(password, salt, secret, associatedData []byte, parallelism, memory, iterations, tagSize uint32) ([]byte, error) {
return DeriveKey(password, salt, secret, associatedData, Params{parallelism, memory, iterations, tagSize})
}
func IDKey(password, salt []byte, time, memory uint32, threads uint8, keyLen uint32) ([]byte, error) {
return DeriveKey(password, salt, nil, nil, Params{uint32(threads), memory, time, keyLen})
}
func (p Params) validate() error {
if p.Iterations < 1 {
return ErrTimeTooSmall
}
if p.Parallelism < 1 {
return ErrLanesTooFew
}
if p.Parallelism > 0xFFFFFF {
return ErrLanesTooMany
}
if uint64(p.Memory) < uint64(8)*uint64(p.Parallelism) {
return ErrMemoryTooSmall
}
if p.TagSize < 4 {
return ErrTagTooShort
}
return nil
}
func initialHash(password, salt, secret, ad []byte, p, m, t, tagSize uint32) [72]byte {
const header = 6*4 + 4*4
total := header + len(password) + len(salt) + len(secret) + len(ad)
var stack [1024]byte
var buf []byte
if total <= len(stack) {
buf = stack[:0]
} else {
buf = make([]byte, 0, total)
}
var w [4]byte
put := func(v uint32) {
binary.LittleEndian.PutUint32(w[:], v)
buf = append(buf, w[:]...)
}
put(p)
put(tagSize)
put(m)
put(t)
put(Version)
put(argon2idVal)
for _, b := range [][]byte{password, salt, secret, ad} {
put(uint32(len(b)))
buf = append(buf, b...)
}
var h0 [72]byte
blake2b(h0[:64], buf)
return h0
}
func initBlocks(mem []Block, h0 *[72]byte, p, laneLength uint32) {
var raw [blockBytes]byte
for lane := uint32(0); lane < p; lane++ {
base := lane * laneLength
for idx := uint32(0); idx < 2; idx++ {
binary.LittleEndian.PutUint32(h0[64:68], idx)
binary.LittleEndian.PutUint32(h0[68:72], lane)
hPrime(raw[:], h0[:])
for i := 0; i < blockWords; i++ {
mem[base+idx][i] = binary.LittleEndian.Uint64(raw[i*8:])
}
}
}
}
func fillSegment(mem []Block, pass, lane, slice, lanes, laneLength, segmentLength, passes, memBlocks uint32) {
dataIndependent := pass == 0 && slice < 2
var addressBlock, inputBlock, zero Block
if dataIndependent {
inputBlock[0] = uint64(pass)
inputBlock[1] = uint64(lane)
inputBlock[2] = uint64(slice)
inputBlock[3] = uint64(memBlocks)
inputBlock[4] = uint64(passes)
inputBlock[5] = argon2idVal
}
index := uint32(0)
if pass == 0 && slice == 0 {
index = 2
if dataIndependent {
inputBlock[6]++
processBlock(&addressBlock, &inputBlock, &zero, false)
processBlock(&addressBlock, &addressBlock, &zero, false)
}
}
offset := lane*laneLength + slice*segmentLength + index
for index < segmentLength {
prev := offset - 1
if index == 0 && slice == 0 {
prev += laneLength
}
var rand uint64
if dataIndependent {
if index%blockWords == 0 {
inputBlock[6]++
processBlock(&addressBlock, &inputBlock, &zero, false)
processBlock(&addressBlock, &addressBlock, &zero, false)
}
rand = addressBlock[index%blockWords]
} else {
rand = mem[prev][0]
}
refLane := uint32(rand>>32) % lanes
if pass == 0 && slice == 0 {
refLane = lane
}
refIndex := indexAlpha(uint32(rand&0xFFFFFFFF), refLane, lane, laneLength, segmentLength, pass, slice, index)
processBlock(&mem[offset], &mem[prev], &mem[refLane*laneLength+refIndex], pass != 0)
index++
offset++
}
}
func indexAlpha(rand, refLane, lane, laneLength, segmentLength, pass, slice, index uint32) uint32 {
sameLane := refLane == lane
var areaSize, startPos uint32
if pass == 0 {
if slice == 0 {
areaSize = index - 1
} else if sameLane {
areaSize = slice*segmentLength + index - 1
} else if index == 0 {
areaSize = slice*segmentLength - 1
} else {
areaSize = slice * segmentLength
}
} else {
if sameLane {
areaSize = laneLength - segmentLength + index - 1
} else if index == 0 {
areaSize = laneLength - segmentLength - 1
} else {
areaSize = laneLength - segmentLength
}
}
x := (uint64(rand) * uint64(rand)) >> 32
rel := uint64(areaSize) - 1 - ((uint64(areaSize) * x) >> 32)
if pass != 0 {
if slice == syncPoints-1 {
startPos = 0
} else {
startPos = (slice + 1) * segmentLength
}
}
return (startPos + uint32(rel)) % laneLength
}
func processBlock(out, in1, in2 *Block, withXor bool) {
var t Block
for i := range t {
t[i] = in1[i] ^ in2[i]
}
permute(&t)
if withXor {
for i := range t {
out[i] ^= in1[i] ^ in2[i] ^ t[i]
}
return
}
for i := range t {
out[i] = in1[i] ^ in2[i] ^ t[i]
}
}
func permute(b *Block) {
for i := 0; i < blockWords; i += 16 {
blamkaRound(b,
i+0, i+1, i+2, i+3, i+4, i+5, i+6, i+7,
i+8, i+9, i+10, i+11, i+12, i+13, i+14, i+15)
}
for i := 0; i < 16; i += 2 {
blamkaRound(b,
i, i+1, 16+i, 16+i+1, 32+i, 32+i+1, 48+i, 48+i+1,
64+i, 64+i+1, 80+i, 80+i+1, 96+i, 96+i+1, 112+i, 112+i+1)
}
}
func blamkaRound(v *Block, i0, i1, i2, i3, i4, i5, i6, i7, i8, i9, i10, i11, i12, i13, i14, i15 int) {
blamkaG(v, i0, i4, i8, i12)
blamkaG(v, i1, i5, i9, i13)
blamkaG(v, i2, i6, i10, i14)
blamkaG(v, i3, i7, i11, i15)
blamkaG(v, i0, i5, i10, i15)
blamkaG(v, i1, i6, i11, i12)
blamkaG(v, i2, i7, i8, i13)
blamkaG(v, i3, i4, i9, i14)
}
func blamkaG(v *Block, a, b, c, d int) {
va, vb, vc, vd := v[a], v[b], v[c], v[d]
va += vb + 2*uint64(uint32(va))*uint64(uint32(vb))
vd = rotr64(vd^va, 32)
vc += vd + 2*uint64(uint32(vc))*uint64(uint32(vd))
vb = rotr64(vb^vc, 24)
va += vb + 2*uint64(uint32(va))*uint64(uint32(vb))
vd = rotr64(vd^va, 16)
vc += vd + 2*uint64(uint32(vc))*uint64(uint32(vd))
vb = rotr64(vb^vc, 63)
v[a], v[b], v[c], v[d] = va, vb, vc, vd
}
func extractTag(mem []Block, p, laneLength, tagSize uint32) []byte {
last := (p-1)*laneLength + laneLength - 1
for lane := uint32(0); lane < p-1; lane++ {
src := mem[lane*laneLength+laneLength-1]
for i := range mem[last] {
mem[last][i] ^= src[i]
}
}
var raw [blockBytes]byte
for i, w := range mem[last] {
binary.LittleEndian.PutUint64(raw[i*8:], w)
}
tag := make([]byte, tagSize)
hPrime(tag, raw[:])
return tag
}
func hPrime(out, in []byte) {
t := uint32(len(out))
var lenPrefix [4]byte
binary.LittleEndian.PutUint32(lenPrefix[:], t)
var stack [4 + 1024]byte
var buf []byte
if 4+len(in) <= len(stack) {
buf = stack[:4+len(in)]
} else {
buf = make([]byte, 4+len(in))
}
copy(buf, lenPrefix[:])
copy(buf[4:], in)
if t <= 64 {
blake2b(out, buf)
return
}
v := blake2b512(buf)
r := (t+31)/32 - 2
off := uint32(0)
for i := uint32(0); i < r; i++ {
copy(out[off:], v[:32])
off += 32
if i == r-1 {
blake2b(out[off:], v[:])
} else {
v = blake2b512(v[:])
}
}
}
export.gopackage argon2id
func Blake2b(out, in []byte) { blake2b(out, in) }
func HPrime(out, in []byte) { hPrime(out, in) }
func CompressBlock(out, x, y *Block) { processBlock(out, x, y, false) }
func Permute(b *Block) { permute(b) }
func Hash(password, salt, secret, associatedData []byte, parallelism, memory, iterations, tagSize uint32) ([]byte, error) {
return Argon2id(password, salt, secret, associatedData, parallelism, memory, iterations, tagSize)
}
The package keeps zero external dependencies; a separate harness (in /tmp/xcheck) drove the differential tests.
cd ~/argon2id
gofmt -l . # prints nothing
go vet ./...
go test -count=1 -v ./...
Result:
--- PASS: TestPeakAllocation64MiB
--- PASS: TestBlake2bVectors
--- PASS: TestRFC9106Vector
--- PASS: TestPublishedVectors
--- PASS: TestMinimumMemory
--- PASS: TestValidation
--- PASS: TestCompressBlockDigest
--- PASS: TestCompressBlockIsPermutationPlusFeedForward
PASS
ok argon2id
p=4, m=32, t=3, password 32×0x01, salt 16×0x02, secret 8×0x03, AD 12×0x04:
0d640df58d78766c08c037a34a8b53c9d01ef0452d75b65eb52520e96b01e659
TestRFC9106Vector reproduces this exactly (exercises p>1, t>1, secret, AD, and H').
(m,t,p) = (64,1,1), (64,2,1), (64,2,2), (256,3,2), (4096,4,4), (1024,4,8), (64,2,3), (1024,3,6) — all 24-byte tags match.
t=1p=1, m=8, t=1, T=4 -> 6b7a947d
p=1, m=8, t=1, T=32 -> f137f8e1...bd63eeb3
p=1, m=8, t=2, T=32 -> fdb4ddb6...b1071646
p=2, m=16, t=1, T=32 -> c1d56bc8...cd249aa7
p=1, m=9,10,11,12 (rounding path)
golang.org/x/crypto/argon2checked 1468 vectors, 0 failures
big-input check: 25 vectors, 0 failures
1468 random (p,m,t,T) combinations with T from 4 to 1024, plus 25 large-input vectors (up to 5000-byte password/salt/secret/AD) against a copy of the reference that exposes secret/AD.
BenchmarkArgon2id64MiB-16 3 202861020 ns/op 67108938 B/op 2 allocs/op
m=65536 allocates the exact 64 MiB matrix (65536 × 1024 = 67,108,864 B) plus the 32-byte tag, in 2 allocations. TestPeakAllocation64MiB asserts the total is within [64 MiB, 64 MiB + 256 KiB].
G is pinned to SHA-256 2002d1681c0b62e7334fb247a714604f4288aa4f4136ecd9f10cdd6897b697f8, and G(X,Y) = P(X xor Y) xor (X xor Y) is checked directly.
t=0, p=0, p>2^24-1, m<8p, T<4, and salt<8 are each rejected with their dedicated error.
tag, err := argon2id.Argon2id(
[]byte("password"), // password
[]byte("somesalt"), // salt (>= 8 bytes)
nil, // secret key (optional)
nil, // associated data (optional)
4, // p, lanes
64*1024, // m, KiB
3, // t, passes
32, // tag length
)
The full self-contained write-up (with all test files embedded) is at ~/argon2id/SOLUTION.md.
# Evidence - Problem class: go-argon2id-rfc9106-block-permutation-vectors - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-27T04:13:08.119Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement Argon2id (RFC 9106) from scratch in Go with no external crypto libraries: Blake2b, the variable-length hash H', the 128-byte block compression function G with its BlaMka multiply-add, the p-lane memory fill using data-independent addressing for the first two slices of the first pass and data-dependent Argon2d-style addressing for the rest, and the final block XOR plus tag hash. The API must accept password, salt, secret key, associated data, parallelism p, memory m in KiB, iterations t, and tag length, and must reject out-of-range parameters. Success = byte-exact equality with the published Argon2id test vectors (including t>1, p>1, and non-empty secret/associated-data variants), peak allocation at m=65536 measured as 64 MiB, and correct behaviour at the 8 KiB minimum memory and t=1 boundary.", "environment": "go1.26", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-argon2id-rfc9106-block-permutation-vectors", "provider": "openrouter", "solved_at": "2026-09-27T04:13:08.129Z", "version": "1.26"}I've implemented and verified the solution. The complete write-up is saved at ~/argon2id/SOLUTION.md, and the runnable package lives in ~/argon2id/. Below is the solution.
Implement Argon2id exactly as specified in RFC 9106 in pure Go with no external crypto libraries: BLAKE2b, the variable-length hash H', the 1024-byte compression G with BlaMka multiply-add, the p-lane memory fill (data-independent addressing for the first two slices of pass 0, Argon2d-style afterwards), and the final XOR + tag hash.
Acceptance criteria: byte-exact RFC vectors (including t>1, p>1, secret key, associated data), peak allocation of 64 MiB at m=65536, correct 8 KiB/t=1 boundary, and rejection of out-of-range parameters.
G: the second message word goes into the second a, not into ca = a + b + x
d = rotr(d ^ a, 32)
c = c + d
b = rotr(b ^ c, 24)
a = a + b + y <-- y is added here
d = rotr(d ^ a, 16)
c = c + d
b = rotr(b ^ c, 63)
The classic mistranscription puts y into c = c + d. The trap: BLAKE2b("") still passes because an all-zero message makes the permutation message-independent, so BLAKE2b("abc") is the first vector that catches it. This was the first real bug found.
indexAlpha: J1² needs a 64-bit productx = (J1 * J1) >> 32
y = (reference_area_size * x) >> 32
z = reference_area_size - 1 - y
If J1 is a uint32, J1*J1 wraps modulo 2³² and >>32 is always 0, collapsing the reference-index distribution. Fix: x := (uint64(rand) * uint64(rand)) >> 32.
H0 binds requested m; the fill uses m' = 4·p·floor(m/(4p))Address blocks embed the rounded m' (the reference's memory_blocks). Mixing these breaks any m not a multiple of 4p.
Data-independent addressing applies only to pass == 0 && slice < 2, from counter blocks G(G(input,0)) where input = LE64(pass)||LE64(lane)||LE64(slice)||LE64(m')||LE64(t)||LE64(type)||LE64(counter). All other segments use the previous block's first word (J1|J2). In pass 0 / slice 0 the lane is forced to the current lane.
B[i][0] = H'^1024(H0||LE32(0)||LE32(i)), B[i][1] = H'^1024(H0||LE32(1)||LE32(i)); pass 0 / slice 0 starts at index 2.
H' for T > 64r = ceil(T/32)-2, output is V1[0:32] || … || Vr[0:32] || H^(T-32r)(Vr).
One []Block (Block = [128]uint64 = 1024 B) is 64 MiB at m=65536; H0/H' scratch is on the stack, so only the matrix and returned tag hit the heap.
go.mod:
module argon2id
go 1.26
blake2b.gopackage argon2id
import "encoding/binary"
var blake2bIV = [8]uint64{
0x6a09e667f3bcc908, 0xbb67ae8584caa73b,
0x3c6ef372fe94f82b, 0xa54ff53a5f1d36f1,
0x510e527fade682d1, 0x9b05688c2b3e6c1f,
0x1f83d9abfb41bd6b, 0x5be0cd19137e2179,
}
var blake2bSigma = [12][16]uint8{
{0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15},
{14, 10, 4, 8, 9, 15, 13, 6, 1, 12, 0, 2, 11, 7, 5, 3},
{11, 8, 12, 0, 5, 2, 15, 13, 10, 14, 3, 6, 7, 1, 9, 4},
{7, 9, 3, 1, 13, 12, 11, 14, 2, 6, 5, 10, 4, 0, 15, 8},
{9, 0, 5, 7, 2, 4, 10, 15, 14, 1, 11, 12, 6, 8, 3, 13},
{2, 12, 6, 10, 0, 11, 8, 3, 4, 13, 7, 5, 15, 14, 1, 9},
{12, 5, 1, 15, 14, 13, 4, 10, 0, 7, 6, 3, 9, 2, 8, 11},
{13, 11, 7, 14, 12, 1, 3, 9, 5, 0, 15, 4, 8, 6, 2, 10},
{6, 15, 14, 9, 11, 3, 0, 8, 12, 2, 13, 7, 1, 4, 10, 5},
{10, 2, 8, 4, 7, 6, 1, 5, 15, 11, 9, 14, 3, 12, 13, 0},
{0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15},
{14, 10, 4, 8, 9, 15, 13, 6, 1, 12, 0, 2, 11, 7, 5, 3},
}
func rotr64(x uint64, n uint) uint64 { return x>>n | x<<(64-n) }
func blake2bCompress(h *[8]uint64, block []byte, t uint64, final bool) {
var m [16]uint64
for i := range m {
m[i] = binary.LittleEndian.Uint64(block[i*8:])
}
var v [16]uint64
copy(v[:8], h[:])
copy(v[8:], blake2bIV[:])
v[12] ^= t
if final {
v[14] = ^v[14]
}
for r := 0; r < 12; r++ {
s := &blake2bSigma[r]
blake2bG(&v, 0, 4, 8, 12, m[s[0]], m[s[1]])
blake2bG(&v, 1, 5, 9, 13, m[s[2]], m[s[3]])
blake2bG(&v, 2, 6, 10, 14, m[s[4]], m[s[5]])
blake2bG(&v, 3, 7, 11, 15, m[s[6]], m[s[7]])
blake2bG(&v, 0, 5, 10, 15, m[s[8]], m[s[9]])
blake2bG(&v, 1, 6, 11, 12, m[s[10]], m[s[11]])
blake2bG(&v, 2, 7, 8, 13, m[s[12]], m[s[13]])
blake2bG(&v, 3, 4, 9, 14, m[s[14]], m[s[15]])
}
for i := 0; i < 8; i++ {
h[i] ^= v[i] ^ v[i+8]
}
}
func blake2bG(v *[16]uint64, a, b, c, d int, x, y uint64) {
v[a] += v[b] + x
v[d] = rotr64(v[d]^v[a], 32)
v[c] += v[d]
v[b] = rotr64(v[b]^v[c], 24)
v[a] += v[b] + y
v[d] = rotr64(v[d]^v[a], 16)
v[c] += v[d]
v[b] = rotr64(v[b]^v[c], 63)
}
func blake2b(out, in []byte) {
if len(out) == 0 || len(out) > 64 {
panic("argon2id: blake2b digest size out of range")
}
h := blake2bIV
h[0] ^= 0x01010000 ^ uint64(len(out))
var t uint64
for len(in) > 128 {
t += 128
blake2bCompress(&h, in[:128], t, false)
in = in[128:]
}
var last [128]byte
copy(last[:], in)
t += uint64(len(in))
blake2bCompress(&h, last[:], t, true)
for i := 0; i < len(out); i++ {
out[i] = byte(h[i>>3] >> (8 * uint(i&7)))
}
}
func blake2b512(in []byte) [64]byte {
var d [64]byte
blake2b(d[:], in)
return d
}
argon2id.go// Package argon2id implements Argon2id per RFC 9106 with only the Go standard library.
package argon2id
import (
"encoding/binary"
"errors"
"math"
)
const Version = 0x13
const (
blockWords = 128
blockBytes = 1024
syncPoints = 4
argon2idVal = 2
)
type Block [blockWords]uint64
type Params struct {
Parallelism uint32 // p, 1 <= p <= 2^24-1
Memory uint32 // m, KiB, >= 8*p
Iterations uint32 // t, >= 1
TagSize uint32 // T, >= 4
}
var (
ErrTimeTooSmall = errors.New("argon2id: iterations t must be >= 1")
ErrMemoryTooSmall = errors.New("argon2id: memory m must be >= 8*parallelism")
ErrLanesTooFew = errors.New("argon2id: parallelism p must be >= 1")
ErrLanesTooMany = errors.New("argon2id: parallelism p must be <= 2^24-1")
ErrTagTooShort = errors.New("argon2id: tag length must be >= 4")
ErrSaltTooShort = errors.New("argon2id: salt must be >= 8 bytes")
)
func DeriveKey(password, salt, secret, associatedData []byte, params Params) ([]byte, error) {
if err := params.validate(); err != nil {
return nil, err
}
if len(salt) < 8 {
return nil, ErrSaltTooShort
}
if uint64(len(password)) > math.MaxUint32 || uint64(len(salt)) > math.MaxUint32 ||
uint64(len(secret)) > math.MaxUint32 || uint64(len(associatedData)) > math.MaxUint32 {
return nil, errors.New("argon2id: input too long")
}
p, m, t, tagSize := params.Parallelism, params.Memory, params.Iterations, params.TagSize
h0 := initialHash(password, salt, secret, associatedData, p, m, t, tagSize)
memBlocks := (m / (syncPoints * p)) * (syncPoints * p) // m' >= 8p
laneLength := memBlocks / p
segmentLength := laneLength / syncPoints
mem := make([]Block, memBlocks)
initBlocks(mem, &h0, p, laneLength)
for pass := uint32(0); pass < t; pass++ {
for slice := uint32(0); slice < syncPoints; slice++ {
for lane := uint32(0); lane < p; lane++ {
fillSegment(mem, pass, lane, slice, p, laneLength, segmentLength, t, memBlocks)
}
}
}
return extractTag(mem, p, laneLength, tagSize), nil
}
func Argon2id(password, salt, secret, associatedData []byte, parallelism, memory, iterations, tagSize uint32) ([]byte, error) {
return DeriveKey(password, salt, secret, associatedData, Params{parallelism, memory, iterations, tagSize})
}
func IDKey(password, salt []byte, time, memory uint32, threads uint8, keyLen uint32) ([]byte, error) {
return DeriveKey(password, salt, nil, nil, Params{uint32(threads), memory, time, keyLen})
}
func (p Params) validate() error {
if p.Iterations < 1 {
return ErrTimeTooSmall
}
if p.Parallelism < 1 {
return ErrLanesTooFew
}
if p.Parallelism > 0xFFFFFF {
return ErrLanesTooMany
}
if uint64(p.Memory) < uint64(8)*uint64(p.Parallelism) {
return ErrMemoryTooSmall
}
if p.TagSize < 4 {
return ErrTagTooShort
}
return nil
}
func initialHash(password, salt, secret, ad []byte, p, m, t, tagSize uint32) [72]byte {
const header = 6*4 + 4*4
total := header + len(password) + len(salt) + len(secret) + len(ad)
var stack [1024]byte
var buf []byte
if total <= len(stack) {
buf = stack[:0]
} else {
buf = make([]byte, 0, total)
}
var w [4]byte
put := func(v uint32) {
binary.LittleEndian.PutUint32(w[:], v)
buf = append(buf, w[:]...)
}
put(p)
put(tagSize)
put(m)
put(t)
put(Version)
put(argon2idVal)
for _, b := range [][]byte{password, salt, secret, ad} {
put(uint32(len(b)))
buf = append(buf, b...)
}
var h0 [72]byte
blake2b(h0[:64], buf)
return h0
}
func initBlocks(mem []Block, h0 *[72]byte, p, laneLength uint32) {
var raw [blockBytes]byte
for lane := uint32(0); lane < p; lane++ {
base := lane * laneLength
for idx := uint32(0); idx < 2; idx++ {
binary.LittleEndian.PutUint32(h0[64:68], idx)
binary.LittleEndian.PutUint32(h0[68:72], lane)
hPrime(raw[:], h0[:])
for i := 0; i < blockWords; i++ {
mem[base+idx][i] = binary.LittleEndian.Uint64(raw[i*8:])
}
}
}
}
func fillSegment(mem []Block, pass, lane, slice, lanes, laneLength, segmentLength, passes, memBlocks uint32) {
dataIndependent := pass == 0 && slice < 2
var addressBlock, inputBlock, zero Block
if dataIndependent {
inputBlock[0] = uint64(pass)
inputBlock[1] = uint64(lane)
inputBlock[2] = uint64(slice)
inputBlock[3] = uint64(memBlocks)
inputBlock[4] = uint64(passes)
inputBlock[5] = argon2idVal
}
index := uint32(0)
if pass == 0 && slice == 0 {
index = 2
if dataIndependent {
inputBlock[6]++
processBlock(&addressBlock, &inputBlock, &zero, false)
processBlock(&addressBlock, &addressBlock, &zero, false)
}
}
offset := lane*laneLength + slice*segmentLength + index
for index < segmentLength {
prev := offset - 1
if index == 0 && slice == 0 {
prev += laneLength
}
var rand uint64
if dataIndependent {
if index%blockWords == 0 {
inputBlock[6]++
processBlock(&addressBlock, &inputBlock, &zero, false)
processBlock(&addressBlock, &addressBlock, &zero, false)
}
rand = addressBlock[index%blockWords]
} else {
rand = mem[prev][0]
}
refLane := uint32(rand>>32) % lanes
if pass == 0 && slice == 0 {
refLane = lane
}
refIndex := indexAlpha(uint32(rand&0xFFFFFFFF), refLane, lane, laneLength, segmentLength, pass, slice, index)
processBlock(&mem[offset], &mem[prev], &mem[refLane*laneLength+refIndex], pass != 0)
index++
offset++
}
}
func indexAlpha(rand, refLane, lane, laneLength, segmentLength, pass, slice, index uint32) uint32 {
sameLane := refLane == lane
var areaSize, startPos uint32
if pass == 0 {
if slice == 0 {
areaSize = index - 1
} else if sameLane {
areaSize = slice*segmentLength + index - 1
} else if index == 0 {
areaSize = slice*segmentLength - 1
} else {
areaSize = slice * segmentLength
}
} else {
if sameLane {
areaSize = laneLength - segmentLength + index - 1
} else if index == 0 {
areaSize = laneLength - segmentLength - 1
} else {
areaSize = laneLength - segmentLength
}
}
x := (uint64(rand) * uint64(rand)) >> 32
rel := uint64(areaSize) - 1 - ((uint64(areaSize) * x) >> 32)
if pass != 0 {
if slice == syncPoints-1 {
startPos = 0
} else {
startPos = (slice + 1) * segmentLength
}
}
return (startPos + uint32(rel)) % laneLength
}
func processBlock(out, in1, in2 *Block, withXor bool) {
var t Block
for i := range t {
t[i] = in1[i] ^ in2[i]
}
permute(&t)
if withXor {
for i := range t {
out[i] ^= in1[i] ^ in2[i] ^ t[i]
}
return
}
for i := range t {
out[i] = in1[i] ^ in2[i] ^ t[i]
}
}
func permute(b *Block) {
for i := 0; i < blockWords; i += 16 {
blamkaRound(b,
i+0, i+1, i+2, i+3, i+4, i+5, i+6, i+7,
i+8, i+9, i+10, i+11, i+12, i+13, i+14, i+15)
}
for i := 0; i < 16; i += 2 {
blamkaRound(b,
i, i+1, 16+i, 16+i+1, 32+i, 32+i+1, 48+i, 48+i+1,
64+i, 64+i+1, 80+i, 80+i+1, 96+i, 96+i+1, 112+i, 112+i+1)
}
}
func blamkaRound(v *Block, i0, i1, i2, i3, i4, i5, i6, i7, i8, i9, i10, i11, i12, i13, i14, i15 int) {
blamkaG(v, i0, i4, i8, i12)
blamkaG(v, i1, i5, i9, i13)
blamkaG(v, i2, i6, i10, i14)
blamkaG(v, i3, i7, i11, i15)
blamkaG(v, i0, i5, i10, i15)
blamkaG(v, i1, i6, i11, i12)
blamkaG(v, i2, i7, i8, i13)
blamkaG(v, i3, i4, i9, i14)
}
func blamkaG(v *Block, a, b, c, d int) {
va, vb, vc, vd := v[a], v[b], v[c], v[d]
va += vb + 2*uint64(uint32(va))*uint64(uint32(vb))
vd = rotr64(vd^va, 32)
vc += vd + 2*uint64(uint32(vc))*uint64(uint32(vd))
vb = rotr64(vb^vc, 24)
va += vb + 2*uint64(uint32(va))*uint64(uint32(vb))
vd = rotr64(vd^va, 16)
vc += vd + 2*uint64(uint32(vc))*uint64(uint32(vd))
vb = rotr64(vb^vc, 63)
v[a], v[b], v[c], v[d] = va, vb, vc, vd
}
func extractTag(mem []Block, p, laneLength, tagSize uint32) []byte {
last := (p-1)*laneLength + laneLength - 1
for lane := uint32(0); lane < p-1; lane++ {
src := mem[lane*laneLength+laneLength-1]
for i := range mem[last] {
mem[last][i] ^= src[i]
}
}
var raw [blockBytes]byte
for i, w := range mem[last] {
binary.LittleEndian.PutUint64(raw[i*8:], w)
}
tag := make([]byte, tagSize)
hPrime(tag, raw[:])
return tag
}
func hPrime(out, in []byte) {
t := uint32(len(out))
var lenPrefix [4]byte
binary.LittleEndian.PutUint32(lenPrefix[:], t)
var stack [4 + 1024]byte
var buf []byte
if 4+len(in) <= len(stack) {
buf = stack[:4+len(in)]
} else {
buf = make([]byte, 4+len(in))
}
copy(buf, lenPrefix[:])
copy(buf[4:], in)
if t <= 64 {
blake2b(out, buf)
return
}
v := blake2b512(buf)
r := (t+31)/32 - 2
off := uint32(0)
for i := uint32(0); i < r; i++ {
copy(out[off:], v[:32])
off += 32
if i == r-1 {
blake2b(out[off:], v[:])
} else {
v = blake2b512(v[:])
}
}
}
export.gopackage argon2id
func Blake2b(out, in []byte) { blake2b(out, in) }
func HPrime(out, in []byte) { hPrime(out, in) }
func CompressBlock(out, x, y *Block) { processBlock(out, x, y, false) }
func Permute(b *Block) { permute(b) }
func Hash(password, salt, secret, associatedData []byte, parallelism, memory, iterations, tagSize uint32) ([]byte, error) {
return Argon2id(password, salt, secret, associatedData, parallelism, memory, iterations, tagSize)
}
The package keeps zero external dependencies; a separate harness (in /tmp/xcheck) drove the differential tests.
cd ~/argon2id
gofmt -l . # prints nothing
go vet ./...
go test -count=1 -v ./...
Result:
--- PASS: TestPeakAllocation64MiB
--- PASS: TestBlake2bVectors
--- PASS: TestRFC9106Vector
--- PASS: TestPublishedVectors
--- PASS: TestMinimumMemory
--- PASS: TestValidation
--- PASS: TestCompressBlockDigest
--- PASS: TestCompressBlockIsPermutationPlusFeedForward
PASS
ok argon2id
p=4, m=32, t=3, password 32×0x01, salt 16×0x02, secret 8×0x03, AD 12×0x04:
0d640df58d78766c08c037a34a8b53c9d01ef0452d75b65eb52520e96b01e659
TestRFC9106Vector reproduces this exactly (exercises p>1, t>1, secret, AD, and H').
(m,t,p) = (64,1,1), (64,2,1), (64,2,2), (256,3,2), (4096,4,4), (1024,4,8), (64,2,3), (1024,3,6) — all 24-byte tags match.
t=1p=1, m=8, t=1, T=4 -> 6b7a947d
p=1, m=8, t=1, T=32 -> f137f8e1...bd63eeb3
p=1, m=8, t=2, T=32 -> fdb4ddb6...b1071646
p=2, m=16, t=1, T=32 -> c1d56bc8...cd249aa7
p=1, m=9,10,11,12 (rounding path)
golang.org/x/crypto/argon2checked 1468 vectors, 0 failures
big-input check: 25 vectors, 0 failures
1468 random (p,m,t,T) combinations with T from 4 to 1024, plus 25 large-input vectors (up to 5000-byte password/salt/secret/AD) against a copy of the reference that exposes secret/AD.
BenchmarkArgon2id64MiB-16 3 202861020 ns/op 67108938 B/op 2 allocs/op
m=65536 allocates the exact 64 MiB matrix (65536 × 1024 = 67,108,864 B) plus the 32-byte tag, in 2 allocations. TestPeakAllocation64MiB asserts the total is within [64 MiB, 64 MiB + 256 KiB].
G is pinned to SHA-256 2002d1681c0b62e7334fb247a714604f4288aa4f4136ecd9f10cdd6897b697f8, and G(X,Y) = P(X xor Y) xor (X xor Y) is checked directly.
t=0, p=0, p>2^24-1, m<8p, T<4, and salt<8 are each rejected with their dedicated error.
tag, err := argon2id.Argon2id(
[]byte("password"), // password
[]byte("somesalt"), // salt (>= 8 bytes)
nil, // secret key (optional)
nil, // associated data (optional)
4, // p, lanes
64*1024, // m, KiB
3, // t, passes
32, // tag length
)
The full self-contained write-up (with all test files embedded) is at ~/argon2id/SOLUTION.md.
# Evidence - Problem class: go-argon2id-rfc9106-block-permutation-vectors - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-27T04:13:08.119Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement Argon2id (RFC 9106) from scratch in Go with no external crypto libraries: Blake2b, the variable-length hash H', the 128-byte block compression function G with its BlaMka multiply-add, the p-lane memory fill using data-independent addressing for the first two slices of the first pass and data-dependent Argon2d-style addressing for the rest, and the final block XOR plus tag hash. The API must accept password, salt, secret key, associated data, parallelism p, memory m in KiB, iterations t, and tag length, and must reject out-of-range parameters. Success = byte-exact equality with the published Argon2id test vectors (including t>1, p>1, and non-empty secret/associated-data variants), peak allocation at m=65536 measured as 64 MiB, and correct behaviour at the 8 KiB minimum memory and t=1 boundary.", "environment": "go1.26", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-argon2id-rfc9106-block-permutation-vectors", "provider": "openrouter", "solved_at": "2026-09-27T04:13:08.129Z", "version": "1.26"}