Goroutine leak merupakan salah satu penyebab utama degradasi performa pada aplikasi Go di lingkungan produksi. Pada framework Go Fiber yang berjalan di atas fasthttp, masalah ini kerap terjadi ketika request context disalahgunakan untuk proses background asynchronous tanpa manajemen lifecycle yang tepat.

Anatomi Masalah: Background Worker Tanpa Terminasi Context

Go Fiber menggunakan fasthttp.RequestCtx di balik layar. Context ini di-pool dan di-recycle oleh runtime setelah response dikembalikan ke client. Menjalankan background goroutine langsung dari handler menggunakan context request Fiber atau channel tanpa mekanisme cancel memicu goroutine yang menggantung permanen (leaked).

Perhatikan contoh handler yang memicu kebocoran berikut:

package main

import (
	"github.com/gofiber/fiber/v2"
)

func RegisterOrderHandler(c *fiber.Ctx) error {
	orderID := c.Params("id")

	// BUG: Goroutine spawn tanpa context cancellation atau channel buffer.
	// Jika worker terhambat, goroutine menumpuk di memori.
	go func(id string) {
		eventCh := make(chan string)
		// Simulasi eksekusi tanpa timeout atau terminasi listener
		go processAuditLog(id, eventCh)
		eventCh <- "ORDER_CREATED" // Tertahan jika receiver tidak membaca
	}(orderID)

	return c.Status(fiber.StatusAccepted).JSON(fiber.Map{
		"status": "processing",
		"id":     orderID,
	})
}

func processAuditLog(id string, ch chan string) {
	// Worker hang tanpa select ke ctx.Done()
	val := <-ch
	_ = val
}

Setiap request HTTP yang masuk akan memicu satu atau lebih goroutine baru yang tertahan pada operasi send/receive channel. Karena alur eksekusi tidak memiliki batas waktu (timeout) dan tidak terikat pada lifecycle application runtime, goroutine ini tidak akan pernah di-garbage collect.

Observasi Prometheus dan Diagnosis Stack Trace via pprof

Deteksi awal kebocoran terlihat dari metrik runtime Go yang dikumpulkan oleh Prometheus. Kueri PromQL berikut memperlihatkan akumulasi goroutine secara linear tanpa penurunan:

go_goroutines{job="fiber-api"}

Gejala penyerta meliputi peningkatan konsumsi memory heap (go_memstats_alloc_bytes) dan penurunan throughput pod. Untuk mendiagnosis titik eksekusi yang macet, pasang middleware pprof pada Fiber:

package main

import (
	"github.com/gofiber/fiber/v2"
	"github.com/gofiber/fiber/v2/middleware/pprof"
)

func SetupRoutes(app *fiber.App) {
	// Expose pprof endpoint untuk profiling runtime
	app.Use(pprof.New())
	
	app.Post("/orders/:id", RegisterOrderHandler)
}

Ambil stack trace goroutine yang sedang berjalan menggunakan command line:

curl -s http://localhost:3000/debug/pprof/goroutine?debug=2 | head -n 30

Output stack trace akan menunjukkan ratusan atau ribuan goroutine dengan status chan send atau chan receive pada fungsi target:

goroutine 8421 [chan send]:
main.RegisterOrderHandler.func1(0xc00018a090, 0x10)
    /app/handlers/order.go:16 +0x68
created by main.RegisterOrderHandler
    /app/handlers/order.go:12 +0x45

goroutine 8422 [chan receive]:
main.processAuditLog(0xc00018a090, 0x10, 0xc00021e060)
    /app/handlers/order.go:27 +0x34
created by main.RegisterOrderHandler.func1
    /app/handlers/order.go:15 +0x55

Trace ini mengonfirmasi bahwa baris 16 dan 27 pada order.go menjadi titik kebocoran.

Prosedur Fast Rollback Produksi

Saat indikator memori mendekati batas OOMKilled pada cluster Kubernetes, jangan tunggu patch kode dibuat. Lakukan rollback segera ke versi image sebelumnya untuk menstabilkan layanan.

Eksekusi perintah rollback pada deployment:

kubectl rollout undo deployment/fiber-api-deployment -n production

Pantau status deployment hingga rolling update selesai:

kubectl rollout status deployment/fiber-api-deployment -n production

Jika pod lama mengonsumsi resource terlalu tinggi dan menahan scheduling pod baru, hapus pod yang bermasalah secara bertahap:

kubectl delete pod -l app=fiber-api --grace-period=30 -n production

Setelah rollback selesai, pastikan grafik go_goroutines di Prometheus kembali ke batas baseline normal (umumnya stabil di angka puluhan hingga ratusan tergantung worker pool aktif).

Postmortem Insiden

KomponenDetail
Waktu Kejadian14:00 - 14:32 UTC
Root CauseHandler RegisterOrderHandler pada rilis v1.2.0 menggunakan unbuffered channel dan sub-goroutine tanpa context cancellation, menyebabkan goroutine tertahan pada runtime park.
DampakPeningkatan memory usage dari 120MB menjadi 1.8GB per pod. Latensi p99 naik dari 25ms ke 950ms karena overhead GC scheduling. 2 pod mengalami restart akibat OOMKilled.
Timeline
  • 14:00: Deployment v1.2.0 selesai via CI/CD.
  • 14:15: Alert HighGoroutineCount aktif (> 5.000 goroutines).
  • 14:22: Tim on-call menganalisis dump pprof dan menemukan bottleneck channel.
  • 14:26: Eksekusi kubectl rollout undo ke v1.1.9.
  • 14:32: Seluruh traffic beralih ke pod lama. Metrik normal.

Mitigasi Teknis: Perbaikan Pola dan Deteksi via goleak di CI

Langkah perbaikan membutuhkan dua tindakan: memperbaiki lifecycle goroutine dan menambahkan assertion test otomatis pada pipeline CI.

1. Refactoring Worker dengan Context dan Timeout

Gunakan context eksplisit yang independen dari request lifecycle Fiber jika proses harus berjalan secara asynchronous:

package main

import (
	"context"
	"time"

	"github.com/gofiber/fiber/v2"
)

func FixedOrderHandler(c *fiber.Ctx) error {
	orderID := c.Params("id")

	// Buat context terpisah dengan timeout eksplisit
	ctx, cancel := context.WithTimeout(context.Background(), 2*time.Second)
	
	go func(ctx context.Context, cancel context.CancelFunc, id string) {
		defer cancel()
		
		eventCh := make(chan string, 1) // Gunakan buffered channel
		go processAuditLogSafe(ctx, id, eventCh)

		select {
		case eventCh <- "ORDER_CREATED":
		case <-ctx.Done():
			// Timeout atau dibatalkan, goroutine keluar
			return
		}
	}(ctx, cancel, orderID)

	return c.Status(fiber.StatusAccepted).JSON(fiber.Map{"status": "accepted"})
}

func processAuditLogSafe(ctx context.Context, id string, ch chan string) {
	select {
	case val := <-ch:
		_ = val
	case <-ctx.Done():
		return
	}
}

2. Implementasi go.uber.org/goleak pada Unit Test

Gunakan library go.uber.org/goleak untuk memverifikasi bahwa tidak ada goroutine yang tertinggal setelah unit test selesai dieksekusi.

Unduh dependency:

go get go.uber.org/goleak

Terapkan pada test file (misalnya order_test.go):

package main

import (
	"net/http/httptest"
	"testing"
	"time"

	"github.com/gofiber/fiber/v2"
	"go.uber.org/goleak"
)

func TestOrderHandler_NoGoroutineLeak(t *testing.T) {
	// Pastikan tidak ada goroutine bocor setelah fungsi test selesai
	defer goleak.VerifyNone(t)

	app := fiber.New()
	app.Post("/orders/:id", FixedOrderHandler)

	req := httptest.NewRequest("POST", "/orders/12345", nil)
	resp, err := app.Test(req)
	if err != nil {
		t.Fatalf("Request failed: %v", err)
	}

	if resp.StatusCode != fiber.StatusAccepted {
		t.Fatalf("Expected status 202, got %d", resp.StatusCode)
	}

	// Beri jeda singkat agar goroutine asynchronous menyelesaikan eksekusi
	time.Sleep(50 * time.Millisecond)
}

3. Otomasi pada Pipeline CI (GitHub Actions)

Tambahkan langkah pengujian goroutine leak ke dalam workflow CI untuk memblokir pull request yang berpotensi menyebabkan regresi:

name: Go Quality & Tests

on:
  pull_request:
    branches: [ main ]
  push:
    branches: [ main ]

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - name: Checkout Code
        uses: actions/checkout@v4

      - name: Set up Go
        uses: actions/setup-go@v5
        with:
          go-version: '1.22'
          cache: true

      - name: Run Leak & Unit Tests
        run: |
          go test -race -v -count=1 ./...

Flag -race melengkapi pengujian dengan mendeteksi kondisi data race, sementara goleak.VerifyNone(t) memastikan seluruh thread runtime Go bersih sebelum binary di-compile dan dikirim ke container registry.