Go Under the Hood · Part 6

Struct Layout and Memory Alignment in Go

Explore Go struct padding, field order, alignment, and size with runnable examples, plus practical guidance on memory savings, portability, and encoding.

Two structs contain the same fields, yet one occupies 24 bytes and the other 16. Nothing was compressed and no information was removed. The difference is the space between fields and at the end of the value.

A struct’s size includes padding needed by its memory layout. Understanding that padding helps explain memory usage, but a smaller struct is not automatically a faster application.

This is Part 6 of Go Under the Hood, following Maps: Semantics Before Internals. You should understand basic structs, arrays, and slices. The experiments use Go 1.27.1 on macOS ARM64 with the standard library only. Their exact sizes are observations for that toolchain and target, not universal Go constants.

Size, alignment, and offset answer different questions

QuantityQuestionInspection tool
SizeHow many bytes does this value occupy, including padding?unsafe.Sizeof(value)
AlignmentAt what address multiples can this value be placed?unsafe.Alignof(value)
Field offsetHow far from the struct’s start does this field begin?unsafe.Offsetof(value.Field)

Sizeof measures the value itself, excluding memory reached through its pointers or descriptors. Alignof(s.Field) reports the field’s alignment requirement within a struct. Offsetof accepts a field selector. These inspection operations do not require converting a pointer or reading arbitrary memory. unsafe package documentation

If a field requires eight-byte alignment, its address must be divisible by eight. The compiler may need a gap before it. The enclosing value also needs a suitable size so successive elements of an array remain aligned.

The language specifies alignment guarantees, including that a struct’s alignment is the largest alignment of its fields, with a minimum of one. Exact layout details depend on the implementation. Go specification: size and alignment

Experiment 1: measure two field orders

Using a Go 1.27 toolchain, create a module:

go version
mkdir layout-demo
cd layout-demo
go mod init example.com/layout-demo
go mod edit -go=1.27.0

Save this as main.go:

package main

import (
	"fmt"
	"runtime"
	"unsafe"
)

type Scattered struct {
	Active bool
	Count  int64
	Ready  bool
}

type Grouped struct {
	Count  int64
	Active bool
	Ready  bool
}

func main() {
	var a Scattered
	var b Grouped
	fmt.Println(runtime.Version(), runtime.GOOS, runtime.GOARCH)
	fmt.Println("scattered size/align:", unsafe.Sizeof(a), unsafe.Alignof(a))
	fmt.Println("scattered offsets:", unsafe.Offsetof(a.Active),
		unsafe.Offsetof(a.Count), unsafe.Offsetof(a.Ready))
	fmt.Println("grouped size/align:", unsafe.Sizeof(b), unsafe.Alignof(b))
	fmt.Println("grouped offsets:", unsafe.Offsetof(b.Count),
		unsafe.Offsetof(b.Active), unsafe.Offsetof(b.Ready))
	fmt.Println("arrays of three:", unsafe.Sizeof([3]Scattered{}),
		unsafe.Sizeof([3]Grouped{}))
}

Run:

go run .
go vet ./...

Output on the verification target:

go1.27.1 darwin arm64
scattered size/align: 24 8
scattered offsets: 0 8 16
grouped size/align: 16 8
grouped offsets: 0 8 9
arrays of three: 72 48

Both types hold ten bytes of field data on this target: eight for the integer and one for each boolean. The rest is padding.

Follow the padding byte by byte

Here are the measured layouts; ranges are byte offsets within each value:

Scattered: 24 bytes
0       Active
1–7     padding
8–15    Count
16      Ready
17–23   trailing padding

Grouped: 16 bytes
0–7     Count
8       Active
9       Ready
10–15   trailing padding

In Scattered, putting Count immediately after Active would start it at offset one. The measured layout moves it to offset eight. After Ready, the total size rounds up to another multiple of eight.

In Grouped, the integer starts at zero and the booleans follow it. There is still trailing padding, but the large internal gap disappears.

Why not stop Grouped at ten bytes? In the measured [3]Grouped layout, each element starts 16 bytes after the previous one. That keeps the next element’s integer aligned. The array output makes this consequence visible without inspecting addresses.

For ordinary fields, grouping those with stronger alignment requirements often reduces gaps. Treat this as a candidate layout to measure, not a universal sorting rule. Nested types and zero-sized fields deserve inspection too; reasoning only from the sum of field sizes is insufficient.

When the difference matters

For one local value, saving eight bytes is rarely a compelling reason to rearrange a clear domain model. For a large collection of inline records, the multiplication can matter.

At the measured sizes, an array of one million records requires 24,000,000 versus 16,000,000 bytes for its elements: an arithmetic difference of 8,000,000 bytes. This is derived from element size, not a measured reduction in process memory. Allocation rounding, other objects, unused slice capacity, and application state still contribute to actual usage.

SituationUseful next stepReason to be cautious
Large slice of compact records retained in memoryInspect element size and candidate field ordersMeasure total workload memory before claiming savings
One request configuration structPrefer clear field groupingSmall layout differences may be immaterial
Struct dominated by strings, slices, or pointersInvestigate referenced data tooShallow size can hide the main cost
Exported type or schema-bound recordReview callers and encoders before reorderingDeclaration order can have observable consequences
Suspected throughput or latency problemBenchmark the affected operationSmaller size alone does not establish a speedup

For a batch-processing service, a useful hypothesis might be: “These retained records dominate memory, and their padding is material.” Verify the record count and retention first. If the real cost is payload buffers, reordering two boolean fields may barely change the result.

This is the same principle used in A Measurement-First Performance Investigation: connect a local change to the observed system problem.

Experiment 2: shallow size is not retained memory

Replace main.go with:

package main

import (
	"fmt"
	"unsafe"
)

type Message struct {
	Label string
	Body  []byte
}

func main() {
	m := Message{Label: "event", Body: make([]byte, 1024)}
	fmt.Println("message:", unsafe.Sizeof(m))
	fmt.Println("body descriptor:", unsafe.Sizeof(m.Body))
	fmt.Println("body len/cap:", len(m.Body), cap(m.Body))
	m.Body = make([]byte, 4096)
	fmt.Println("after replacement:", unsafe.Sizeof(m), len(m.Body))
}

Output on the same target:

message: 40
body descriptor: 24
body len/cap: 1024 1024
after replacement: 40 4096

The larger body does not change the struct’s size. Its field still contains a slice descriptor; the backing bytes live separately. Similarly, the string field’s size does not include its text. This is the direct-versus-referenced-memory distinction documented by Sizeof. unsafe.Sizeof

A common mistake is multiplying Sizeof(Message{}) by the number of messages and calling that total retained memory. It omits bodies and labels, and shared backing storage makes naive recursive addition unreliable too. Use a heap investigation when the question concerns live application memory.

As Part 3 showed, a small view can retain a larger backing array. Field order cannot fix that ownership or retention issue.

Memory layout is not an encoding format

Do not use a struct’s raw memory as a portable file or network format. A format needs explicit field representation and byte order; Go’s in-memory padding does not define those rules.

Replace main.go with:

package main

import (
	"bytes"
	"encoding/binary"
	"fmt"
	"unsafe"
)

type Record struct {
	Kind  uint8
	Count uint64
	Flags uint8
}

func main() {
	r := Record{Kind: 1, Count: 2, Flags: 3}
	var out bytes.Buffer
	if err := binary.Write(&out, binary.LittleEndian, r); err != nil {
		panic(err)
	}
	fmt.Println("memory/encoded:", unsafe.Sizeof(r), out.Len())
	fmt.Printf("encoded bytes: % x\n", out.Bytes())
}

Output:

memory/encoded: 24 10
encoded bytes: 01 02 00 00 00 00 00 00 00 03

encoding/binary writes successive fixed-size fields using the selected byte order. Implicit compiler padding is not emitted. Explicit blank fields can represent format padding when needed. This API requires fixed-size values or supported slices of them; arbitrary strings and slice-bearing structs need another encoding design. encoding/binary.Write

Here the wire representation is ten bytes even though the struct occupies 24. Reordering the Go fields would change the order written by this encoder. If a protocol already exists, an internal layout improvement must not silently change its format. A separate wire type or explicit field encoding can preserve the contract.

Keep portability and API compatibility in the decision

The measured offsets above are not promises for every CPU target. Even with fixed-width integer fields, alignment can differ; int, uintptr, and pointer sizes add further target dependence. Record go version and the target architecture whenever reporting layouts. Go specification: numeric types and alignment

Run the experiment on each relevant target. Cross-compiling verifies that a program builds for another target; it does not mean that the foreign executable has been run. Avoid a test requiring “this struct must be 16 bytes everywhere” unless your supported platforms and implementation contract justify it.

Before reordering fields, inspect composite literals and consumers. Keyed literals associate values with field names. Positional literals associate values with declaration order: reordering same-typed fields can silently swap their meaning, while other changes cause compilation errors. Go specification: composite literals

For internal records, prefer keyed construction and check any reflective or serialisation code. For public types, weigh the compatibility cost against a measured benefit. A layout tweak is not worth an accidental change to a customer’s encoded data.

Exercise: predict a different layout

In Experiment 1, replace both int64 fields with int32. On the verified ARM64 target, predict 12 bytes for Scattered and eight for Grouped, each aligned to four. The arrays of three become 36 and 24 bytes. Run the program and explain every gap.

Next, move Ready beside Active in Scattered, before Count. With the original int64, that ordering also occupies 16 bytes on this target. The important change is eliminating a gap, not mechanically putting every largest field first.

Finally, in Experiment 3 move Flags before Count. The in-memory size falls to 16, but the encoded bytes begin 01 03 02. Explain why an optimisation that helps an internal collection could break a protocol consumer.

The next part, Methods, Interfaces, and Typed Nil, moves from physical layout to behaviour: how receivers and interface values affect method calls and nil checks.

Sources

Back to the journal