# Profiling และ Benchmarking Go 2026: pprof, trace และคำถามสัมภาษณ์ > คู่มือ profiling Go ด้วย pprof และ runtime/trace อย่างครอบคลุม เทคนิควิเคราะห์ CPU, หน่วยความจำ และ goroutine เพื่อเพิ่มประสิทธิภาพและเตรียมตัวสัมภาษณ์ - Published: 2026-09-19 - Updated: 2026-09-19 - Author: Anthony Fillion-Maillet - Reading time: 5 min --- การ profiling Go ด้วย pprof และ package `runtime/trace` เปลี่ยนการคาดเดาเป็นข้อมูลที่ชัดเจน คำถามเกี่ยวกับประสิทธิภาพปรากฏในการสัมภาษณ์ Go เกือบทุกครั้ง และผู้สมัครที่สามารถอ่าน flame graph หรืออธิบายว่าเมื่อใดควรใช้ `-inuse_space` เทียบกับ `-allocs` จะโดดเด่น เครื่องมือนี้มาพร้อมกับ standard library ไม่ต้องพึ่งพา dependency ภายนอก และผสานรวมกับ benchmark โดยตรง > **ประเภท Profile ที่ควรรู้** > > Go 1.24+ มี profile ในตัว 7 แบบ: CPU, heap, allocs, goroutine, threadcreate, block และ mutex Go 1.26 เพิ่ม profile goroutine leak ทดลองที่ตรวจจับ goroutine ที่ถูก block และเข้าถึงไม่ได้ ## CPU Profiling ด้วย pprof: จุดเริ่มต้น CPU profiling สุ่มตัวอย่าง call stack ในช่วงเวลาที่กำหนด (ค่าเริ่มต้น 100 Hz) และบันทึกว่าฟังก์ชันใดใช้เวลาประมวลผล Package `runtime/pprof` จัดการการเก็บข้อมูลระดับต่ำ ในขณะที่ `go tool pprof` วิเคราะห์ผลลัพธ์ Profile 30 วินาทีเก็บ 3000 ตัวอย่าง ซึ่งเพียงพอสำหรับนัยสำคัญทางสถิติในแอปพลิเคชันส่วนใหญ่ โปรแกรม standalone เปิดใช้ profiling โดยเรียก `pprof.StartCPUProfile` ตอนเริ่มต้น Profile ถูกเขียนลงไฟล์ที่ `go tool pprof` อ่านภายหลัง: ```go // main.go package main import ( "flag" "log" "os" "runtime/pprof" ) var cpuprofile = flag.String("cpuprofile", "", "write cpu profile to file") func main() { flag.Parse() if *cpuprofile != "" { f, err := os.Create(*cpuprofile) if err != nil { log.Fatal(err) } defer f.Close() pprof.StartCPUProfile(f) defer pprof.StopCPUProfile() } // Application logic ที่นี่ } ``` สำหรับ HTTP server ให้ import `net/http/pprof` เป็น side effect Package นี้ลงทะเบียน handler ที่ `/debug/pprof/` โดยอัตโนมัติ ไม่ต้องเปลี่ยนโค้ดนอกจาก import: ```go // server.go package main import ( "net/http" _ "net/http/pprof" // ลงทะเบียน handler /debug/pprof/* ) func main() { http.HandleFunc("/", handler) http.ListenAndServe(":8080", nil) } ``` ดึง CPU profile 30 วินาทีจาก server ที่กำลังทำงานด้วย `go tool pprof http://localhost:8080/debug/pprof/profile?seconds=30` เครื่องมือจะดาวน์โหลด profile และเปิด shell แบบโต้ตอบ พารามิเตอร์ `seconds` ควบคุมระยะเวลาการเก็บข้อมูล ## วิเคราะห์ Profile: top, list และ Flame Graph Shell ของ pprof มีคำสั่งเพื่อระบุ bottleneck `top` แสดงฟังก์ชันที่ใช้ CPU มากที่สุด เรียงตาม flat time คำสั่ง `list` แสดง source code พร้อมคำอธิบาย timing ต่อบรรทัด ชี้ไปที่บรรทัดที่ครองการทำงาน ```bash # เซสชัน terminal กับ go tool pprof $ go tool pprof cpu.prof (pprof) top 10 Showing nodes accounting for 4.2s, 85% of 4.9s total flat flat% sum% cum cum% 1.8s 36.73% 36.73% 1.8s 36.73% runtime.memmove 0.9s 18.37% 55.10% 0.9s 18.37% encoding/json.(*decodeState).scanWhile 0.5s 10.20% 65.30% 2.3s 46.94% main.processRecords ... (pprof) list processRecords Total: 4.9s 0.5s 2.3s (flat, cum) 46.94% of Total 20: for _, r := range records { 21: 0.3s 0.3s data := json.Marshal(r) 22: 0.2s 2.0s result := transform(data) ... ``` อินเทอร์เฟซเว็บเพิ่มการวิเคราะห์แบบภาพ รัน `go tool pprof -http=:6060 cpu.prof` เพื่อเปิดเบราว์เซอร์พร้อม flame graph, directed graph และ source view ตั้งแต่ Go 1.26 flame graph ปรากฏเป็นมุมมองเริ่มต้นใน UI เว็บ Flame graph แสดงลำดับชั้นการเรียกในแนวนอน โดยแถบที่กว้างกว่าแสดงว่าใช้เวลามากขึ้นในฟังก์ชันนั้นและฟังก์ชันที่เรียก > **การอ่าน Flame Graph** > > ใน flame graph แกน x แสดงประชากรของตัวอย่าง ไม่ใช่เวลา แต่ละกล่องคือฟังก์ชัน และความกว้างแสดงว่าฟังก์ชันนั้นปรากฏในตัวอย่างบ่อยแค่ไหน ฟังก์ชันแม่อยู่ใต้ฟังก์ชันลูก มองหา plateau กว้างที่ด้านบน: ฟังก์ชันเหล่านั้นทำงานจริง ## Memory Profiling: Heap vs Allocs Memory profiling ตอบคำถามสองข้อที่แตกต่างกัน Profile **heap** (`-inuse_space`) แสดงว่าอะไรเก็บหน่วยความจำไว้ในขณะที่เก็บข้อมูล Profile **allocs** แสดงว่าการจัดสรรเกิดขึ้นที่ไหนตลอดเวลา แม้ว่าหน่วยความจำนั้นจะถูกปล่อยแล้ว เพื่อลดการใช้หน่วยความจำปัจจุบัน ตรวจสอบ heap profile เพื่อลดอัตราการจัดสรรและแรงกดดัน GC ตรวจสอบ allocs profile อัตราการจัดสรรสูงทำให้เกิด garbage collection บ่อย ซึ่งหยุด goroutine และเพิ่มการใช้ CPU ```bash # ดึง heap profile จาก server ที่กำลังทำงาน $ curl -o heap.prof http://localhost:8080/debug/pprof/heap $ go tool pprof -inuse_space heap.prof # ดึง allocs profile (จำนวนการจัดสรรใน 30 วินาที) $ curl -o allocs.prof "http://localhost:8080/debug/pprof/allocs?seconds=30" $ go tool pprof -alloc_objects allocs.prof ``` Flag `-inuse_objects` นับอ็อบเจกต์ที่ยังอยู่แทนไบต์ มีประโยชน์สำหรับการระบุ memory fragmentation Flag `-alloc_space` แสดงไบต์ทั้งหมดที่จัดสรรในช่วง profile เผยให้เห็นฟังก์ชันที่ประมวลผลหน่วยความจำแม้ว่าจะปล่อยอย่างรวดเร็ว จุดร้อนการจัดสรรทั่วไปรวมถึงการต่อ string ใน loop (ใช้ `strings.Builder`) การแปลง interface ที่ escape ไป heap และการขยาย slice โดยไม่มี preallocate [เอกสาร compiler ของ Go](https://go.dev/doc/diagnostics#profiling) อธิบาย escape analysis อย่างละเอียด ## Benchmark Profiling ด้วย testing.B Package `testing` ผสานรวม profiling โดยตรงกับ benchmark การรวมกันนี้แยก code path เฉพาะโดยไม่มี noise จากแอปพลิเคชันเต็มรูปแบบ Benchmark profiling ตอบคำถาม: "ฟังก์ชันนี้ทำงานอย่างไรเมื่อแยกตัว?" ```go // parser_test.go package parser import "testing" func BenchmarkParseJSON(b *testing.B) { data := []byte(`{"id":1,"name":"test","values":[1,2,3]}`) b.ReportAllocs() // รวมสถิติการจัดสรร b.ResetTimer() // ไม่รวม setup ใน timing for i := 0; i < b.N; i++ { _, _ = Parse(data) } } ``` สร้าง profile ระหว่างการรัน benchmark ด้วย flag Flag `-cpuprofile` และ `-memprofile` เขียน profile ลงไฟล์เพื่อวิเคราะห์ภายหลัง: ```bash # CPU profile ระหว่าง benchmark $ go test -bench=BenchmarkParseJSON -cpuprofile=cpu.prof -benchtime=5s # Memory profile ระหว่าง benchmark $ go test -bench=BenchmarkParseJSON -memprofile=mem.prof -benchtime=5s # วิเคราะห์ผลลัพธ์ $ go tool pprof -http=:6060 cpu.prof ``` Flag `-benchtime` ควบคุมว่า benchmark รันนานเท่าใด รันนานกว่าให้ profile ที่แม่นยำกว่าแต่ใช้เวลามากกว่า รัน 5 วินาทีมักให้ผลลัพธ์ที่เสถียร สำหรับ micro-benchmark ใช้ `-count=10` เพื่อรันหลายรอบและตรวจสอบความแปรปรวน ## Execution Tracing ด้วย runtime/trace ในขณะที่ pprof แสดงว่าเวลาถูกใช้ที่ไหน `runtime/trace` แสดงว่าเหตุการณ์เกิดขึ้นเมื่อใด Trace จับการจัดตาราง goroutine, system call, เหตุการณ์ GC และกิจกรรมเครือข่ายบน timeline ความสามารถในการมองเห็นพฤติกรรม concurrency นี้เสริม profiling ทางสถิติ ```go // trace_example.go package main import ( "os" "runtime/trace" ) func main() { f, _ := os.Create("trace.out") defer f.Close() trace.Start(f) defer trace.Stop() // Application logic runConcurrentTasks() } ``` Trace viewer แสดงช่วงชีวิต goroutine, เหตุการณ์ blocking และการใช้งาน processor แต่ละ goroutine ปรากฏเป็นแถบแนวนอน โดยสีแสดงว่ากำลังรัน, ถูก block หรือกำลังรอการจัดตาราง: ```bash $ go tool trace trace.out # เปิดเบราว์เซอร์ที่ http://127.0.0.1:port ``` สำหรับ HTTP server ดึง trace จาก `/debug/pprof/trace?seconds=5` Trace viewer แสดงว่า goroutine ใดถูก block ที่ไหน เผยรูปแบบ contention ที่ CPU profile พลาด มุมมอง "Goroutine analysis" จัดกลุ่ม goroutine ตามตำแหน่งที่สร้าง ช่วยระบุ leak หรือ fan-out ที่ไม่คาดคิด Trace หนักกว่า profile Trace 5 วินาทีของ server ที่ยุ่งอาจสร้างข้อมูลหลายร้อยเมกะไบต์ ใช้ระยะเวลาสั้นและการเก็บที่มีเป้าหมาย [บล็อก Go เกี่ยวกับ execution tracing](https://go.dev/blog/execution-traces-2024) ครอบคลุมเทคนิคการวิเคราะห์ขั้นสูง ## Block และ Mutex Profiling สำหรับ Contention Block profiling บันทึก goroutine ที่กำลังรอบน synchronization primitive: channel, mutex และ condition variable Mutex profiling เน้นเฉพาะ mutex contention Profile เหล่านี้เผย bottleneck ของ concurrency ที่ CPU profiling มองไม่เห็น เปิดใช้ profile เหล่านี้โดยตั้งค่าพารามิเตอร์ runtime ก่อนที่ contention จะเกิด: ```go // เปิดใช้ block profiling (1 = สุ่มตัวอย่างทุกเหตุการณ์ blocking) runtime.SetBlockProfileRate(1) // เปิดใช้ mutex profiling (1 = สุ่มตัวอย่างทุก mutex contention) runtime.SetMutexProfileFraction(1) ``` สำหรับ production ตั้งค่าสูงกว่าเพื่อลด overhead Block profile rate 1000000 (หนึ่งไมโครวินาที) หรือ mutex fraction 100 ให้ข้อมูลที่มีประโยชน์พร้อมผลกระทบน้อยที่สุด การตั้งค่าเหล่านี้ต่ำเกินไปจะจับทุกเหตุการณ์และอาจทำให้แอปพลิเคชันช้าลง ดึง profile เหล่านี้จาก endpoint มาตรฐาน: ```bash $ curl -o block.prof http://localhost:8080/debug/pprof/block $ curl -o mutex.prof http://localhost:8080/debug/pprof/mutex $ go tool pprof block.prof ``` Block profile แสดงเวลารอทั้งหมด ไม่ใช่จำนวนเหตุการณ์ blocking ฟังก์ชันที่ block 1 วินาทีหนึ่งครั้งดูเหมือนกับฟังก์ชันที่ block 1 มิลลิวินาที 1000 ครั้ง ใช้ execution tracing เพื่อแยกแยะกรณีเหล่านี้ ## ข้อผิดพลาดทั่วไปใน Profiling Profiling นำมาซึ่ง overhead ที่อาจทำให้ผลลัพธ์เบี่ยงเบน CPU profiling เพิ่มประมาณ 5% overhead Memory profiling สุ่มตัวอย่างการจัดสรร (1 ต่อ 512KB โดยค่าเริ่มต้น) ดังนั้นการจัดสรรขนาดเล็กอาจไม่ปรากฏ Tracing จับทุกเหตุการณ์และอาจเพิ่ม 10-30% overhead ข้อผิดพลาดหลายอย่างนำไปสู่ profile ที่ทำให้เข้าใจผิด: **Profiling optimized build ต่างกัน** ควร profile ด้วย build flag เดียวกับที่ใช้ใน production เสมอ Debug build ปิด inlining และ optimization ทำให้จุดร้อนปรากฏในที่ต่างกัน **Profiling ภายใต้ load ปลอม** Profile ของ server ที่ว่างแสดง idle loop ไม่ใช่ bottleneck จริง ควร profile ภายใต้รูปแบบ traffic ที่สมจริง **ละเลย overhead GC** CPU profile รวมเวลาที่ใช้ใน garbage collection การปรากฏของ `runtime.gc*` สูงบ่งบอกปัญหาการจัดสรรหน่วยความจำ ไม่ใช่ปัญหา CPU แก้ไขด้วย memory profiling **ระยะเวลา profiling สั้น** Profile 1 วินาทีจับได้เพียง 100 ตัวอย่าง noise ทางสถิติครอบงำ ควร profile อย่างน้อย 30 วินาทีภายใต้ load ที่คงที่ ## คำถามสัมภาษณ์ Go เกี่ยวกับ Profiling คำถามประสิทธิภาพทดสอบว่าผู้สมัครสามารถวินิจฉัยปัญหาจริงได้หรือไม่ ผู้สัมภาษณ์มองหาความคุ้นเคยกับเครื่องมือและความเข้าใจว่าแต่ละ profile เผยอะไร **ถ: เมื่อใดที่ heap profiling แสดงผลลัพธ์ต่างจาก allocs profiling?** Heap แสดงหน่วยความจำที่ถูกเก็บไว้ในขณะที่เก็บข้อมูล Allocs แสดงการจัดสรรทั้งหมด รวมถึงหน่วยความจำที่ปล่อยแล้ว ฟังก์ชันที่จัดสรร buffer ชั่วคราวใน loop ปรากฏใน allocs แต่ไม่ใน heap หาก buffer ถูกเก็บก่อน snapshot ใช้ allocs เพื่อลดแรงกดดัน GC ใช้ heap เพื่อหา leak **ถ: goroutine ดูเหมือนติด Profile ใดช่วยได้?** Goroutine profile แสดง stack trace ของทุก goroutine Block profile แสดงว่า goroutine กำลังรอที่ไหน สำหรับ Go 1.26+ profile goroutine leak ทดลองตรวจจับ goroutine ที่เข้าถึงไม่ได้ที่ถูก block บน channel หรือ mutex Execution tracing แสดง timeline ของเหตุการณ์ blocking **ถ: flat percentage เทียบกับ cumulative percentage ใน pprof หมายถึงอะไร?** Flat วัดเวลาในฟังก์ชันนั้นเอง Cumulative รวมเวลาในฟังก์ชันที่มันเรียก ฟังก์ชันที่มี cumulative สูงแต่ flat ต่ำเป็น coordinator ที่มอบหมายงาน ฟังก์ชันที่มี flat time สูงทำการคำนวณจริง เพิ่มประสิทธิภาพฟังก์ชันที่มี flat time สูงก่อน **ถ: จะ profile benchmark โดยไม่ profile setup ของ test อย่างไร?** เรียก `b.ResetTimer()` หลังจาก setup เสร็จ สำหรับ benchmark ที่มี setup ต่อการวนซ้ำ ใช้ `b.StopTimer()` และ `b.StartTimer()` รอบโค้ด setup การเรียก timer มี overhead ระดับนาโนวินาที ดังนั้นหลีกเลี่ยงใน loop ที่แคบ **ถ: ทำไมฟังก์ชันอาจไม่ปรากฏใน CPU profile แม้ว่าจะช้า?** CPU profiling จับเฉพาะฟังก์ชันที่กำลังใช้ CPU อย่างแข็งขัน ฟังก์ชัน I/O-bound (รอเครือข่าย, disk หรือ channel) ปรากฏใน block profile หรือ trace ไม่ใช่ CPU profile Sampling อาจพลาดฟังก์ชันที่รันรวมน้อยกว่า 10ms **ถ: การใช้งาน Swiss Tables map ของ Go 1.24 ส่งผลต่อ profiling อย่างไร?** Go 1.24 แทนที่ map ที่ใช้ bucket ด้วย Swiss Tables ลด CPU overhead 2-3% สำหรับ workload ที่ใช้ map หนัก Profile ที่เก็บก่อนและหลังอัปเกรดแสดง call stack เกี่ยวกับ map ที่แตกต่างกัน Flag `GOEXPERIMENT=noswissmap` กลับไปใช้การใช้งานเก่าเพื่อเปรียบเทียบ ## Continuous Profiling ใน Production Profile แบบ point-in-time พลาดปัญหาชั่วคราว เครื่องมือ continuous profiling เช่น [Pyroscope](https://pyroscope.io/) หรือ [Parca](https://www.parca.dev/) เก็บตัวอย่าง low-overhead อย่างต่อเนื่อง ทำให้สามารถเปรียบเทียบระหว่าง deployment เครื่องมือเหล่านี้เชื่อมโยง profile กับ metric และ trace Profile ในตัวของ Go ทำงานกับเครื่องมือเหล่านี้ผ่านรูปแบบ pprof บริการ [pprof.me](https://pprof.me/) เพิ่มฟีเจอร์เปรียบเทียบในปี 2026 ทำให้สามารถอัปโหลดและ diff profile เพื่อวัดผลกระทบของการเพิ่มประสิทธิภาพก่อนและหลังเปลี่ยนโค้ด สำหรับ[การเตรียมตัวสัมภาษณ์ Go](/technologies/go/interview-questions/testing) การเข้าใจทั้งเครื่องมือและแนวคิดพื้นฐานมีความสำคัญ [Package context](/technologies/go/interview-questions/context-package) และ [รูปแบบ concurrency](/technologies/go/interview-questions/concurrency-patterns) มักปรากฏพร้อมกับคำถาม profiling การเพิ่มประสิทธิภาพมักต้องรวมข้อมูล profiling กับความรู้เกี่ยวกับพฤติกรรม runtime ของ Go ## Go Profiling เผยอะไรเกี่ยวกับพฤติกรรมแอปพลิเคชัน - CPU profile ระบุฟังก์ชันร้อนแต่พลาดงาน I/O-bound รวมกับ trace เพื่อภาพรวมที่สมบูรณ์ - Memory profile แยกแยะปัญหาการเก็บ (heap) จาก churn การจัดสรร (allocs) ใช้ `-inuse_space` สำหรับ leak ใช้ `-alloc_space` สำหรับแรงกดดัน GC - Block และ mutex profile เผย contention ที่ CPU profile มองไม่เห็น เปิดใช้เมื่อ latency พุ่งสูงภายใต้ load - Benchmark profiling แยก code path เฉพาะ เรียก `b.ReportAllocs()` และ `b.ResetTimer()` เสมอเพื่อการวัดที่แม่นยำ - Trace viewer แสดงการจัดตาราง goroutine และเหตุการณ์ blocking บน timeline จำเป็นสำหรับการวินิจฉัย bug ของ concurrency - การปรับปรุง runtime ของ Go 1.24 ลด CPU overhead 2-3% ผ่าน Swiss Tables map และการใช้งาน mutex ใหม่ - Profiling ใน production ด้วย endpoint `/debug/pprof/` ต้องมีการยืนยันตัวตน อย่าเปิดเผย endpoint เหล่านี้ต่อสาธารณะ: มันเปิดเผยสถานะแอปพลิเคชันภายในและอาจเปิดใช้ denial-of-service ผ่านการเก็บ profile ที่มีราคาแพง --- Source: SharpSkill (https://sharpskill.dev), tech interview preparation for your real stack. HTML version of this page: https://sharpskill.dev/th/blog/go/go-profiling-benchmarking-pprof-trace-interview