local-ai-bench

Qwen3.8-Flash-Next and GLM-5.3-Flash on Ryzen 9 5950X · 128 GB DDR4-3200 · 3× RTX 3090

Overview

152 result records in results, latest run 2026-10-05 00:18 AEST. Every number on this site comes from those records. Failed and crashed runs are listed alongside the passing ones.

Qwen3.8-Flash-Next

running

59 records, no qualified eligible config yet

Target: lab decode ≥ 50 tok/s and lab prefill > 1000 tok/s on the same config, passing §5 stability; kit band top-1 ≥ 0.987, KL ≤ 0.0013.

No qualified headline yet. Best so far

Lab decode
57.4 tok/s ✓
Lab prefill
2,537 tok/s ✓
Stability
failed ✕

Config sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 · current · x8 · EXL3 3.05 bpw

59 runs in 14 configs: ✓ pass 45 ! fail 8 ✕ crash 6 · leaderboard

GLM-5.3-Flash

running

89 records, no qualified eligible config yet

Target: 1× RTX 3090, registry speed gate 15 tok/s (lab C1 decode), §5 stability and quality floor (≥ 80% of the reference solve rate). 3-card configs are a separate recipe.

Best 1 card

Lab decode
10.0 tok/s ✕
Lab prefill
151 tok/s
Quality
not measured
Stability
failed ✕

Config ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 · ik_llama.cpp · GGUF IQ2_M · current · x8 · below 15 tok/s: documented experiment, no registry PR

Best 3 cards

Lab decode
14.0 tok/s ✕
Lab prefill
189 tok/s
Quality
not measured
Stability
failed ✕

Config ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 · ik_llama.cpp · GGUF IQ2_M · current · 3× x8 · below 15 tok/s: documented experiment, no registry PR

89 runs in 22 configs: ✓ pass 39 ! fail 43 ✕ crash 7 · leaderboard

Pull requests

PRs open in Phase 3, after the operator review of this site.

Recent runs

StartedOutcomeModelKindConfigKey metricsStageTicket
2026-10-05 00:18 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2lab_decode_c1_tps 55.8 tok/s · lab_prefill_tps 2,419 tok/sT13
2026-10-05 00:14 AEST✓ passFlash-Nextloadsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2load_seconds 257 sT13
2026-10-05 00:07 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2lab_decode_c1_tps 51.3 tok/s · lab_prefill_tps 2,596 tok/sT14b
2026-10-04 23:57 AEST! failFlash-Nextstabilitysglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2decode_window_min_tps 40.9 tok/s · tool_calls 72 callsT14b
2026-10-04 23:55 AEST✓ passFlash-NextTTFTsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2ttft_prefill_tps 2,022 tok/sT14b
2026-10-04 23:52 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2lab_decode_c1_tps 53.7 tok/s · lab_prefill_tps 2,440 tok/sT14b
2026-10-04 23:48 AEST✓ passFlash-Nextloadsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2load_seconds 215 sT14b
2026-10-04 23:43 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2lab_decode_c1_tps 54.6 tok/s · lab_prefill_tps 2,580 tok/sT14a
2026-10-04 23:33 AEST! failFlash-Nextstabilitysglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2decode_window_min_tps 40.1 tok/s · tool_calls 57 callsT14a
2026-10-04 23:31 AEST✓ passFlash-NextTTFTsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2ttft_prefill_tps 2,014 tok/sT14a

64 failed or crashed runs: see Runs and dropped configs.

Qwen3.8-Flash-Next leaderboard

One row per config_id, aggregating all of its records. Decode/prefill are the best value from a passing record (from a failed one when none passed, marked † and never counted as meeting a target). ✓ = meets the target (decode ≥ 50 tok/s, prefill > 1000 tok/s). Chips link to each run; ! = fail, ✕ = crash, – = skipped.

ConfigLab decode tok/sLab prefill tok/sTTFT prefill tok/sKit C1 tok/sOccupied maxCtx configuredQualityStabilityTool callsEngineQuantCardsLayout · PCIeEligibilityRuns
sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 ✕ 1 crash ! 1 fail57.4 ✓2,537 ✓2,00947.4167k205kin band top-1 0.98967 · KL 0.00097327failed 2 runs · min win 39.771 ✕sglangEXL3 3.05 bpw1current · x8noneload load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel
llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 ! 1 fail57.1 ✓608 ✕987–106k131k–failed 1 run · min win 38.670 ✕llama.cppGGUF IQ4_XS3current · 3× x8noneload lab lab TTFT stability !
sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 ! 1 fail55.6 ✓2,596 ✓2,01748.2167k205kin band top-1 0.98867 · KL 0.00096647failed 1 run · min win 37.967 ✕sglangEXL3 3.05 bpw1current · x8noneload lab lab kit sweep TTFT stability ! quality panel
sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 ! 1 fail55.6 ✓2,580 ✓2,014–167k205k–failed 1 run · min win 40.157 ✕sglangEXL3 3.05 bpw1current · x8noneload lab lab TTFT stability !
fn-trellis-3.05-0xsero-main-x8-gpu2 ! 1 fail55.3 ✓2,228 ✓2,26346.9167k205kin band top-1 0.99234 · KL 0.0009667failed 1 run · min win 38.163 ✕sglangEXL3 3.05 bpw1current · x8noneload lab lab kit sweep TTFT stability ! quality panel
sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 ! 1 fail53.7 ✓2,596 ✓2,022–167k205k–failed 1 run · min win 40.972 ✕sglangEXL3 3.05 bpw1current · x8noneload lab lab TTFT stability !
llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu01244.3 ✕717 ✕1,067–106k131k–––llama.cppGGUF IQ4_XS3current · 3× x8noneload lab TTFT
llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu01235.0 ✕713 ✕980–106k131k–––llama.cppGGUF IQ4_XS3current · 3× x8noneload lab TTFT
fn-llamacpp-iq4xs-128k-1x-fit-gpu2 ! 1 fail26.5 ✕157 ✕269–106k131k–failed 1 run · min win 23.849 ✕llama.cppGGUF IQ4_XS1current · x8noneload lab lab TTFT stability !
fn-llamacpp-iq4xs-128k-1x-gpu2 ✕ 1 crash–––––131k–––llama.cppGGUF IQ4_XS1current · x8noneload ✕
llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012 ! 1 fail–––––131k–failed 1 run0 ✕llama.cppGGUF IQ4_XS3current · 3× x8noneload stability !
sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2 ✕ 1 crash–––––205k–––sglangEXL3 2.05 bpw1current · x8noneload ✕
sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2 ✕ 1 crash–––––205k–––sglangEXL3 4.05 bpw1current · x8noneload ✕
sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2 ✕ 2 crash–––––205k–––sglangEXL3 3.05 bpw1current · x8noneload lab ✕ TTFT ✕

GLM-5.3-Flash leaderboard

One row per config_id, aggregating all of its records. Decode/prefill are the best value from a passing record (from a failed one when none passed, marked † and never counted as meeting a target). ✓ = meets the target (decode ≥ 15 tok/s). Chips link to each run; ! = fail, ✕ = crash, – = skipped.

ConfigLab decode tok/sLab prefill tok/sTTFT prefill tok/sKit C1 tok/sOccupied maxCtx configuredQualityStabilityTool callsEngineQuantCardsLayout · PCIeEligibilityRuns
ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 ! 3 fail14.0 †189 †283–93k131k–failed 1 run · min win 13.30 ✕ik_llama.cppGGUF IQ2_M3current · 3× x8noneload lab ! lab ! TTFT stability !
llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 ✕ 1 crash ! 2 fail12.3 †158 †166–93k131k–failed 1 run · min win 11.426 ✕llama.cppGGUF IQ2_M3current · 3× x8noneload lab ! lab ✕ TTFT stability !
llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 ! 1 fail12.2 †153 †159–93k131k–––llama.cppGGUF IQ2_M3current · 3× x8noneload lab ! TTFT
llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 ! 3 fail11.8 †219 †240–93k131k–failed 1 run · min win 11.024 ✕llama.cppGGUF IQ2_M3current · 3× x8noneload lab ! lab ! TTFT stability !
llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012 ! 1 fail11.2 †258 †298–93k131k–––llama.cppGGUF IQ2_M3current · 3× x8noneload lab ! TTFT
llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 ! 3 fail10.5 †132 †130–93k131k–failed 1 run · min win 9.8622 ✕llama.cppGGUF IQ3_XXS3current · 3× x8noneload lab ! lab ! TTFT stability !
ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 ✕ 1 crash ! 1 fail10.2 †166 †248–93k131k–failed 1 run0 ✕ik_llama.cppGGUF IQ2_M3current · 3× x8noneload lab ! TTFT stability ✕
ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 ! 3 fail10.0 †151 †203–93k131k–failed 1 run · min win 9.460 ✕ik_llama.cppGGUF IQ2_M1current · x8noneload lab ! lab ! TTFT stability !
glm-llamacpp-iq2m-128k-1x-gpu2 ! 3 fail9.60 †81 †86.3–93k131k–failed 1 run · min win 8.9124 ✕llama.cppGGUF IQ2_M1current · x8noneload lab ! lab ! TTFT stability !
llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 ! 3 fail9.60 †288 †321–93k131k–failed 1 run · min win 8.8527 ✕llama.cppGGUF IQ2_M1current · x8noneload lab ! lab ! TTFT stability !
llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 ! 3 fail9.50 †200 †212–93k131k–failed 1 run · min win 8.8530 ✕llama.cppGGUF IQ2_M1current · x8noneload lab ! lab ! TTFT stability !
llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 ! 1 fail9.50 †201 †213–93k131k–––llama.cppGGUF IQ2_M1current · x8noneload lab ! TTFT
llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 ! 3 fail9.30 †80 †84.4–93k131k–failed 1 run · min win 8.8019 ✕llama.cppGGUF Q2_K1current · x8noneload lab ! lab ! TTFT stability !
llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 ! 1 fail9.10 †370 †384–93k131k–––llama.cppGGUF IQ2_M1current · x8noneload lab ! TTFT
ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 ✕ 1 crash ! 1 fail8.60 †147 †199–93k131k–failed 1 run0 ✕ik_llama.cppGGUF IQ2_M1current · x8noneload lab ! TTFT stability ✕
llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 ! 3 fail8.40 †71 †75.9–93k131k–failed 1 run · min win 7.6525 ✕llama.cppGGUF IQ3_XXS1current · x8noneload lab ! lab ! TTFT stability !
llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 ! 3 fail8.30 †156 †149–93k131k–failed 1 run · min win 7.8125 ✕llama.cppGGUF IQ4_XS3current · 3× x8noneload lab ! lab ! TTFT stability !
ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 ✕ 1 crash ! 1 fail8.20 †147 †201–93k131k–failed 1 run0 ✕ik_llama.cppGGUF IQ2_M1current · x8noneload lab ! TTFT stability ✕
exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 ! 4 fail7.50 †739 †483–93k131k–failed 1 run · min win 3.3621 ✕exllamav3EXL3 3.05 bpw1current · x8noneload load lab ! lab ! TTFT ! TTFT stability !
ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012 ✕ 1 crash–––––131k–––ik_llama.cppGGUF IQ2_M3current · 3× x8noneload ✕
ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012 ✕ 1 crash–––––131k–––ik_llama.cppGGUF IQ2_M3current · 3× x8noneload ✕
ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012 ✕ 1 crash–––––131k–––ik_llama.cppGGUF IQ2_M3current · 3× x8noneload ✕

Charts

Each chart appears once records contain its measurements. Hover or tap a mark for its run; every chart has a table view.

Decode vs occupied context

Lab C1 decode and the stability minimum 60 s window against the occupied prompt tokens reported by the server.

Qwen3.8-Flash-NextGLM-5.3-Flashlab C1 (first 30 s)stability: min 60 s window
020406080050100150200occupied context (k tokens)decode tok/sFlash-Next target 50GLM gate 15128kfn-trellis-3.05-0xsero-main-x8-gpu2 · lab C1 decode 49.3 tok/s at 166,667 occupied · passfn-trellis-3.05-0xsero-main-x8-gpu2 · stability min 60 s window 38.1 tok/s at 30,034 occupied · failfn-trellis-3.05-0xsero-main-x8-gpu2 · lab C1 decode 55.3 tok/s at 166,667 occupied · passfn-llamacpp-iq4xs-128k-1x-fit-gpu2 · lab C1 decode 26.5 tok/s at 106,294 occupied · passfn-llamacpp-iq4xs-128k-1x-fit-gpu2 · stability min 60 s window 23.8 tok/s at 16,819 occupied · failfn-llamacpp-iq4xs-128k-1x-fit-gpu2 · lab C1 decode 26.1 tok/s at 106,294 occupied · passglm-llamacpp-iq2m-128k-1x-gpu2 · lab C1 decode 9.60 tok/s at 93,437 occupied · failglm-llamacpp-iq2m-128k-1x-gpu2 · stability min 60 s window 8.91 tok/s at 9,047 occupied · failglm-llamacpp-iq2m-128k-1x-gpu2 · lab C1 decode 9.60 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 · lab C1 decode 9.20 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 · stability min 60 s window 8.80 tok/s at 6,658 occupied · failllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 · lab C1 decode 9.30 tok/s at 93,437 occupied · failik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 · lab C1 decode 10.0 tok/s at 93,437 occupied · failik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 · stability min 60 s window 9.46 tok/s at 1,874 occupied · failik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 · lab C1 decode 9.50 tok/s at 93,437 occupied · failik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 · lab C1 decode 8.20 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 · lab C1 decode 8.30 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 · stability min 60 s window 7.65 tok/s at 4,875 occupied · failllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 · lab C1 decode 8.40 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 · lab C1 decode 9.50 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 · stability min 60 s window 8.85 tok/s at 9,732 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 · lab C1 decode 9.50 tok/s at 93,437 occupied · failik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 · lab C1 decode 8.60 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 · lab C1 decode 12.3 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 · stability min 60 s window 11.4 tok/s at 8,118 occupied · failllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 · lab C1 decode 10.5 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 · stability min 60 s window 9.86 tok/s at 5,309 occupied · failllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 · lab C1 decode 10.4 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 · lab C1 decode 12.2 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 · lab C1 decode 11.8 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 · stability min 60 s window 11.0 tok/s at 6,174 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012 · lab C1 decode 11.8 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 · lab C1 decode 9.50 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 · lab C1 decode 9.60 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 · stability min 60 s window 8.85 tok/s at 6,632 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 · lab C1 decode 9.60 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 · lab C1 decode 8.00 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 · stability min 60 s window 7.81 tok/s at 6,852 occupied · failllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012 · lab C1 decode 8.30 tok/s at 93,437 occupied · failllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 · lab C1 decode 56.9 tok/s at 106,294 occupied · passllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 · stability min 60 s window 38.6 tok/s at 39,226 occupied · failllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012 · lab C1 decode 57.1 tok/s at 106,294 occupied · passllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 · lab C1 decode 9.10 tok/s at 93,437 occupied · failllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012 · lab C1 decode 11.2 tok/s at 93,437 occupied · failik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 · lab C1 decode 14.0 tok/s at 93,437 occupied · failik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 · stability min 60 s window 13.3 tok/s at 1,874 occupied · failik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 · lab C1 decode 13.4 tok/s at 93,437 occupied · failllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012 · lab C1 decode 44.3 tok/s at 106,294 occupied · passllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012 · lab C1 decode 35.0 tok/s at 106,294 occupied · passik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 · lab C1 decode 10.2 tok/s at 93,437 occupied · failsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 · lab C1 decode 55.6 tok/s at 166,667 occupied · passsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 · stability min 60 s window 37.9 tok/s at 36,327 occupied · failsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 · lab C1 decode 54.5 tok/s at 166,667 occupied · passexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 · lab C1 decode 7.40 tok/s at 93,437 occupied · failexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 · stability min 60 s window 3.36 tok/s at 4,589 occupied · failexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 · lab C1 decode 7.50 tok/s at 93,437 occupied · failsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 · lab C1 decode 55.8 tok/s at 166,667 occupied · passsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 · stability min 60 s window 39.7 tok/s at 25,309 occupied · failsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 · lab C1 decode 57.4 tok/s at 166,667 occupied · passsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 · lab C1 decode 55.6 tok/s at 166,667 occupied · passsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 · stability min 60 s window 40.1 tok/s at 29,027 occupied · failsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2 · lab C1 decode 54.6 tok/s at 166,667 occupied · passsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 · lab C1 decode 53.7 tok/s at 166,667 occupied · passsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 · stability min 60 s window 40.9 tok/s at 39,281 occupied · failsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2 · lab C1 decode 51.3 tok/s at 166,667 occupied · pass
Table view
ModelConfigMeasureOccupiedtok/sOutcome
Qwen3.8-Flash-Nextfn-trellis-3.05-0xsero-main-x8-gpu2lab C1 decode166,66749.3✓ pass
Qwen3.8-Flash-Nextfn-trellis-3.05-0xsero-main-x8-gpu2stability min 60 s window30,03438.1! fail
Qwen3.8-Flash-Nextfn-trellis-3.05-0xsero-main-x8-gpu2lab C1 decode166,66755.3✓ pass
Qwen3.8-Flash-Nextfn-llamacpp-iq4xs-128k-1x-fit-gpu2lab C1 decode106,29426.5✓ pass
Qwen3.8-Flash-Nextfn-llamacpp-iq4xs-128k-1x-fit-gpu2stability min 60 s window16,81923.8! fail
Qwen3.8-Flash-Nextfn-llamacpp-iq4xs-128k-1x-fit-gpu2lab C1 decode106,29426.1✓ pass
GLM-5.3-Flashglm-llamacpp-iq2m-128k-1x-gpu2lab C1 decode93,4379.60! fail
GLM-5.3-Flashglm-llamacpp-iq2m-128k-1x-gpu2stability min 60 s window9,0478.91! fail
GLM-5.3-Flashglm-llamacpp-iq2m-128k-1x-gpu2lab C1 decode93,4379.60! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2lab C1 decode93,4379.20! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2stability min 60 s window6,6588.80! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2lab C1 decode93,4379.30! fail
GLM-5.3-Flashik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2lab C1 decode93,43710.0! fail
GLM-5.3-Flashik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2stability min 60 s window1,8749.46! fail
GLM-5.3-Flashik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2lab C1 decode93,4379.50! fail
GLM-5.3-Flashik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2lab C1 decode93,4378.20! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2lab C1 decode93,4378.30! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2stability min 60 s window4,8757.65! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2lab C1 decode93,4378.40! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2lab C1 decode93,4379.50! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2stability min 60 s window9,7328.85! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2lab C1 decode93,4379.50! fail
GLM-5.3-Flashik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2lab C1 decode93,4378.60! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012lab C1 decode93,43712.3! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012stability min 60 s window8,11811.4! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012lab C1 decode93,43710.5! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012stability min 60 s window5,3099.86! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012lab C1 decode93,43710.4! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012lab C1 decode93,43712.2! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012lab C1 decode93,43711.8! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012stability min 60 s window6,17411.0! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012lab C1 decode93,43711.8! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2lab C1 decode93,4379.50! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2lab C1 decode93,4379.60! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2stability min 60 s window6,6328.85! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2lab C1 decode93,4379.60! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012lab C1 decode93,4378.00! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012stability min 60 s window6,8527.81! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012lab C1 decode93,4378.30! fail
Qwen3.8-Flash-Nextllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012lab C1 decode106,29456.9✓ pass
Qwen3.8-Flash-Nextllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012stability min 60 s window39,22638.6! fail
Qwen3.8-Flash-Nextllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012lab C1 decode106,29457.1✓ pass
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2lab C1 decode93,4379.10! fail
GLM-5.3-Flashllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012lab C1 decode93,43711.2! fail
GLM-5.3-Flashik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012lab C1 decode93,43714.0! fail
GLM-5.3-Flashik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012stability min 60 s window1,87413.3! fail
GLM-5.3-Flashik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012lab C1 decode93,43713.4! fail
Qwen3.8-Flash-Nextllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012lab C1 decode106,29444.3✓ pass
Qwen3.8-Flash-Nextllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012lab C1 decode106,29435.0✓ pass
GLM-5.3-Flashik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012lab C1 decode93,43710.2! fail
Qwen3.8-Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2lab C1 decode166,66755.6✓ pass
Qwen3.8-Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2stability min 60 s window36,32737.9! fail
Qwen3.8-Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2lab C1 decode166,66754.5✓ pass
GLM-5.3-Flashexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2lab C1 decode93,4377.40! fail
GLM-5.3-Flashexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2stability min 60 s window4,5893.36! fail
GLM-5.3-Flashexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2lab C1 decode93,4377.50! fail
Qwen3.8-Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2lab C1 decode166,66755.8✓ pass
Qwen3.8-Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2stability min 60 s window25,30939.7! fail
Qwen3.8-Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2lab C1 decode166,66757.4✓ pass
Qwen3.8-Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2lab C1 decode166,66755.6✓ pass
Qwen3.8-Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2stability min 60 s window29,02740.1! fail
Qwen3.8-Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2lab C1 decode166,66754.6✓ pass
Qwen3.8-Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2lab C1 decode166,66753.7✓ pass
Qwen3.8-Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2stability min 60 s window39,28140.9! fail
Qwen3.8-Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2lab C1 decode166,66751.3✓ pass

Speed vs quality

Best lab decode per config against its quality measurement.

Qwen3.8-Flash-Next: decode vs mean KL (lower KL is better)

02040608000.00050.0010.0015mean KL vs exllamav3 reference panellab decode tok/starget 50band KL ≤ 0.0013fn-trellis-3.05-0xsero-main-x8-gpu2 · KL 0.0009667 · top-1 0.99234 · decode 55.3fn-trellis-3.05-0xsero-main-x8-gpu2sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2 · KL 0.00096647 · top-1 0.98867 · decode 55.6sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2 · KL 0.00097327 · top-1 0.98967 · decode 57.4sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Table view
ConfigKLtop-1Lab decode
fn-trellis-3.05-0xsero-main-x8-gpu20.00096670.9923455.3
sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu20.000966470.9886755.6
sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu20.000973270.9896757.4

1 vs 3 cards (current layout)

Best config per card count, same model.

Best lab decode

020406080Flash-Next EXL3 3.05 bpw · 1 cardsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2: 57.4 tok/s decode57.4Flash-Next GGUF IQ4_XS · 3 cardsllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012: 57.1 tok/s decode57.1GLM GGUF IQ2_M · 1 cardik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2: 10.0 tok/s decode10.0GLM GGUF IQ2_M · 3 cardsik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012: 14.0 tok/s decode14.0tok/s

Lab prefill of the same configs

01k2k3kFlash-Next EXL3 3.05 bpw · 1 cardsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2: 2,537 tok/s prefill2,537Flash-Next GGUF IQ4_XS · 3 cardsllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012: 608 tok/s prefill608GLM GGUF IQ2_M · 1 cardik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2: 151 tok/s prefill151GLM GGUF IQ2_M · 3 cardsik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012: 189 tok/s prefill189tok/s
Table view

x8 vs x16

No measurements yet: needs single-card lab results in both layouts for a model.

All runs

Every record, newest first, failed, crashed and skipped runs included.

StartedOutcomeModelKindConfigKey metricsStageTicket
2026-10-05 00:18 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2lab_decode_c1_tps 55.8 tok/s · lab_prefill_tps 2,419 tok/sT13
2026-10-05 00:14 AEST✓ passFlash-Nextloadsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2load_seconds 257 sT13
2026-10-05 00:07 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2lab_decode_c1_tps 51.3 tok/s · lab_prefill_tps 2,596 tok/sT14b
2026-10-04 23:57 AEST! failFlash-Nextstabilitysglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2decode_window_min_tps 40.9 tok/s · tool_calls 72 callsT14b
2026-10-04 23:55 AEST✓ passFlash-NextTTFTsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2ttft_prefill_tps 2,022 tok/sT14b
2026-10-04 23:52 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2lab_decode_c1_tps 53.7 tok/s · lab_prefill_tps 2,440 tok/sT14b
2026-10-04 23:48 AEST✓ passFlash-Nextloadsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2load_seconds 215 sT14b
2026-10-04 23:43 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2lab_decode_c1_tps 54.6 tok/s · lab_prefill_tps 2,580 tok/sT14a
2026-10-04 23:33 AEST! failFlash-Nextstabilitysglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2decode_window_min_tps 40.1 tok/s · tool_calls 57 callsT14a
2026-10-04 23:31 AEST✓ passFlash-NextTTFTsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2ttft_prefill_tps 2,014 tok/sT14a
2026-10-04 23:26 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2lab_decode_c1_tps 55.6 tok/s · lab_prefill_tps 2,430 tok/sT14a
2026-10-04 23:22 AEST✓ passFlash-Nextloadsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2load_seconds 254 sT14a
2026-10-04 23:13 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2lab_decode_c1_tps 57.4 tok/s · lab_prefill_tps 2,537 tok/sT14a
2026-10-04 23:02 AEST! failFlash-Nextstabilitysglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2decode_window_min_tps 39.7 tok/s · tool_calls 71 callsT14a
2026-10-04 22:58 AEST✓ passFlash-Nextloadsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2load_seconds 264 sT14a
2026-10-04 22:54 AEST✓ passGLMTTFTexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2ttft_prefill_tps 483 tok/sT23
2026-10-04 22:49 AEST✓ passGLMloadexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2load_seconds 276 sT23
2026-10-04 22:48 AEST✕ crashFlash-NextTTFTsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2ttft_prefill_tps unavailablemeasureT14a
2026-10-04 22:47 AEST✕ crashFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2lab_decode_c1_tps unavailable · lab_prefill_tps unavailableharnessT14a
2026-10-04 22:42 AEST✓ passFlash-Nextloadsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2load_seconds 257 sT14a
2026-10-04 22:32 AEST✕ crashFlash-Nextstabilitysglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2decode_window_min_tps unavailable · tool_calls 0 callsmeasureT13
2026-10-04 22:32 AEST✓ passFlash-Nextquality panelsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2top1_agreement 0.98967 fraction · mean_kl 0.00097327 natsT13
2026-10-04 22:30 AEST✓ passFlash-NextTTFTsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2ttft_prefill_tps 2,009 tok/sT13
2026-10-04 22:15 AEST✓ passFlash-Nextkit sweepsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2decode_c1_tps 47.4 tok/s · prefill_32k_tps 2,678 tok/sT13
2026-10-04 22:10 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2lab_decode_c1_tps 55.8 tok/s · lab_prefill_tps 2,419 tok/sT13
2026-10-04 22:05 AEST✓ passFlash-Nextloadsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2load_seconds 260 sT13
2026-10-04 21:27 AEST! failGLMlabexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2lab_decode_c1_tps 7.50 tok/s · lab_prefill_tps 720 tok/sgatesT23
2026-10-04 21:17 AEST! failGLMstabilityexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2decode_window_min_tps 3.36 tok/s · tool_calls 21 callsT23
2026-10-04 21:13 AEST! failGLMTTFTexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2ttft_prefill_tps unavailableinvalid measurementT23
2026-10-04 21:02 AEST! failGLMlabexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2lab_decode_c1_tps 7.40 tok/s · lab_prefill_tps 739 tok/sgatesT23
2026-10-04 20:57 AEST✓ passGLMloadexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2load_seconds 300 sT23
2026-10-04 20:56 AEST✕ crashFlash-Nextloadsglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2load_seconds unavailableloadT15
2026-10-04 20:52 AEST✕ crashFlash-Nextloadsglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2load_seconds unavailableloadT15
2026-10-04 20:47 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2lab_decode_c1_tps 54.5 tok/s · lab_prefill_tps 2,596 tok/sT13
2026-10-04 20:36 AEST! failFlash-Nextstabilitysglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2decode_window_min_tps 37.9 tok/s · tool_calls 67 callsT13
2026-10-04 20:36 AEST✓ passFlash-Nextquality panelsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2top1_agreement 0.98867 fraction · mean_kl 0.00096647 natsT13
2026-10-04 20:34 AEST✓ passFlash-NextTTFTsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2ttft_prefill_tps 2,017 tok/sT13
2026-10-04 20:19 AEST✓ passFlash-Nextkit sweepsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2decode_c1_tps 48.2 tok/s · prefill_32k_tps 2,678 tok/sT13
2026-10-04 20:14 AEST✓ passFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2lab_decode_c1_tps 55.6 tok/s · lab_prefill_tps 2,433 tok/sT13
2026-10-04 20:10 AEST✓ passFlash-Nextloadsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2load_seconds 236 sT13
2026-10-04 19:38 AEST✕ crashGLMstabilityik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012decode_window_min_tps unavailable · tool_calls 0 callsmeasureT22c
2026-10-04 19:26 AEST✓ passGLMTTFTik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012ttft_prefill_tps 248 tok/sT22c
2026-10-04 19:12 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012lab_decode_c1_tps 10.2 tok/s · lab_prefill_tps 166 tok/sgatesT22c
2026-10-04 19:11 AEST✓ passGLMloadik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012load_seconds 48.6 sT22c
2026-10-04 19:11 AEST✓ passnonehw bwhw-bw-gpu1-currenth2d_gbps 13.5 GB/s · d2h_gbps 13.2 GB/sT05a
2026-10-04 18:19 AEST✓ passFlash-NextTTFTllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012ttft_prefill_tps 980 tok/sT11
2026-10-04 18:11 AEST✓ passFlash-Nextlabllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012lab_decode_c1_tps 35.0 tok/s · lab_prefill_tps 713 tok/sT11
2026-10-04 18:10 AEST✓ passFlash-Nextloadllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012load_seconds 39.9 sT11
2026-10-04 18:06 AEST✓ passFlash-NextTTFTllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012ttft_prefill_tps 1,067 tok/sT11
2026-10-04 18:01 AEST✓ passFlash-Nextlabllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012lab_decode_c1_tps 44.3 tok/s · lab_prefill_tps 717 tok/sT11
2026-10-04 18:00 AEST✓ passFlash-Nextloadllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012load_seconds 39.9 sT11
2026-10-04 17:42 AEST✕ crashGLMloadik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012load_seconds unavailableloadT22c
2026-10-04 17:31 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012lab_decode_c1_tps 13.4 tok/s · lab_prefill_tps 189 tok/sgatesT22c
2026-10-04 17:21 AEST! failGLMstabilityik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012decode_window_min_tps 13.3 tok/s · tool_calls 0 callsT22c
2026-10-04 17:10 AEST✓ passGLMTTFTik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012ttft_prefill_tps 283 tok/sT22c
2026-10-04 16:58 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012lab_decode_c1_tps 14.0 tok/s · lab_prefill_tps 185 tok/sgatesT22c
2026-10-04 16:58 AEST✓ passGLMloadik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012load_seconds 44.3 sT22c
2026-10-04 16:48 AEST✓ passGLMTTFTllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012ttft_prefill_tps 298 tok/sT21
2026-10-04 16:30 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012lab_decode_c1_tps 11.2 tok/s · lab_prefill_tps 258 tok/sgatesT21
2026-10-04 16:30 AEST✓ passGLMloadllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012load_seconds 54.9 sT21
2026-10-04 16:23 AEST✓ passGLMTTFTllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2ttft_prefill_tps 384 tok/sT20b
2026-10-04 16:14 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2lab_decode_c1_tps 9.10 tok/s · lab_prefill_tps 370 tok/sgatesT20b
2026-10-04 16:13 AEST✓ passGLMloadllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2load_seconds 48.9 sT20b
2026-10-04 16:00 AEST✓ passFlash-Nextlabllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012lab_decode_c1_tps 57.1 tok/s · lab_prefill_tps 608 tok/sT11
2026-10-04 15:49 AEST! failFlash-Nextstabilityllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012decode_window_min_tps 38.6 tok/s · tool_calls 70 callsT11
2026-10-04 15:45 AEST✓ passFlash-NextTTFTllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012ttft_prefill_tps 987 tok/sT11
2026-10-04 15:39 AEST✓ passFlash-Nextlabllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012lab_decode_c1_tps 56.9 tok/s · lab_prefill_tps 605 tok/sT11
2026-10-04 15:38 AEST✓ passFlash-Nextloadllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012load_seconds 39.9 sT11
2026-10-04 15:19 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012lab_decode_c1_tps 8.30 tok/s · lab_prefill_tps 156 tok/sgatesT21
2026-10-04 15:08 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012decode_window_min_tps 7.81 tok/s · tool_calls 25 callsT21
2026-10-04 14:52 AEST✓ passGLMTTFTllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012ttft_prefill_tps 149 tok/sT21
2026-10-04 14:26 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012lab_decode_c1_tps 8.00 tok/s · lab_prefill_tps 139 tok/sgatesT21
2026-10-04 14:25 AEST✓ passGLMloadllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012load_seconds 91.2 sT21
2026-10-04 14:24 AEST✕ crashGLMloadik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012load_seconds unavailableloadT22c
2026-10-04 14:23 AEST✕ crashGLMloadik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012load_seconds unavailableloadT22c
2026-10-04 14:14 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2lab_decode_c1_tps 9.60 tok/s · lab_prefill_tps 288 tok/sgatesT20b
2026-10-04 14:04 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2decode_window_min_tps 8.85 tok/s · tool_calls 27 callsT20b
2026-10-04 13:55 AEST✓ passGLMTTFTllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2ttft_prefill_tps 321 tok/sT20b
2026-10-04 13:42 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2lab_decode_c1_tps 9.60 tok/s · lab_prefill_tps 288 tok/sgatesT20b
2026-10-04 13:41 AEST✓ passGLMloadllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2load_seconds 45.9 sT20b
2026-10-04 13:29 AEST✓ passGLMTTFTllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2ttft_prefill_tps 213 tok/sT20b
2026-10-04 13:06 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2lab_decode_c1_tps 9.50 tok/s · lab_prefill_tps 201 tok/sgatesT20b
2026-10-04 13:05 AEST✓ passGLMloadllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2load_seconds 45.9 sT20b
2026-10-04 12:49 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012lab_decode_c1_tps 11.8 tok/s · lab_prefill_tps 218 tok/sgatesT21
2026-10-04 12:39 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012decode_window_min_tps 11.0 tok/s · tool_calls 24 callsT21
2026-10-04 12:27 AEST✓ passGLMTTFTllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012ttft_prefill_tps 240 tok/sT21
2026-10-04 12:16 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012lab_decode_c1_tps 11.8 tok/s · lab_prefill_tps 219 tok/sgatesT21
2026-10-04 12:15 AEST✓ passGLMloadllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012load_seconds 54.9 sT21
2026-10-04 11:59 AEST✓ passGLMTTFTllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012ttft_prefill_tps 159 tok/sT21
2026-10-04 11:39 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012lab_decode_c1_tps 12.2 tok/s · lab_prefill_tps 153 tok/sgatesT21
2026-10-04 11:38 AEST✓ passGLMloadllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012load_seconds 55.0 sT21
2026-10-04 10:42 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012lab_decode_c1_tps 10.4 tok/s · lab_prefill_tps 132 tok/sgatesT21
2026-10-04 10:31 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012decode_window_min_tps 9.86 tok/s · tool_calls 22 callsT21
2026-10-04 10:12 AEST✓ passGLMTTFTllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012ttft_prefill_tps 130 tok/sT21
2026-10-04 09:30 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012lab_decode_c1_tps 10.5 tok/s · lab_prefill_tps 126 tok/sgatesT21
2026-10-04 09:29 AEST✓ passGLMloadllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012load_seconds 79.7 sT21
2026-10-04 09:17 AEST✕ crashGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012lab_decode_c1_tps unavailable · lab_prefill_tps unavailableharnessT21
2026-10-04 09:07 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012decode_window_min_tps 11.4 tok/s · tool_calls 26 callsT21
2026-10-04 08:51 AEST✓ passGLMTTFTllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012ttft_prefill_tps 166 tok/sT21
2026-10-04 08:27 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012lab_decode_c1_tps 12.3 tok/s · lab_prefill_tps 158 tok/sgatesT21
2026-10-04 08:25 AEST✓ passGLMloadllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012load_seconds 101 sT21
2026-10-04 08:17 AEST! failFlash-Nextstabilityllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012decode_window_min_tps unavailable · tool_calls 0 callsT11
2026-10-04 08:06 AEST✓ passFlash-Nextloadllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012load_seconds 39.9 sT11
2026-10-04 06:02 AEST✕ crashGLMstabilityik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2decode_window_min_tps unavailable · tool_calls 0 callsmeasureT22b
2026-10-04 05:48 AEST✓ passGLMTTFTik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2ttft_prefill_tps 199 tok/sT22b
2026-10-04 05:33 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2lab_decode_c1_tps 8.60 tok/s · lab_prefill_tps 147 tok/sgatesT22b
2026-10-04 05:32 AEST✓ passGLMloadik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2load_seconds 51.9 sT22b
2026-10-04 05:16 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2lab_decode_c1_tps 9.50 tok/s · lab_prefill_tps 199 tok/sgatesT20b
2026-10-04 05:06 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2decode_window_min_tps 8.85 tok/s · tool_calls 30 callsT20b
2026-10-04 04:54 AEST✓ passGLMTTFTllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2ttft_prefill_tps 212 tok/sT20b
2026-10-04 04:26 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2lab_decode_c1_tps 9.50 tok/s · lab_prefill_tps 200 tok/sgatesT20b
2026-10-04 04:26 AEST✓ passGLMloadllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2load_seconds 45.9 sT20b
2026-10-04 03:40 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2lab_decode_c1_tps 8.40 tok/s · lab_prefill_tps 71 tok/sgatesT20b
2026-10-04 03:30 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2decode_window_min_tps 7.65 tok/s · tool_calls 25 callsT20b
2026-10-04 02:55 AEST✓ passGLMTTFTllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2ttft_prefill_tps 75.9 tok/sT20b
2026-10-04 02:26 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2lab_decode_c1_tps 8.30 tok/s · lab_prefill_tps 65 tok/sgatesT20b
2026-10-04 02:25 AEST✓ passGLMloadllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2load_seconds 58.2 sT20b
2026-10-04 01:51 AEST✕ crashGLMstabilityik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2decode_window_min_tps unavailable · tool_calls 0 callsmeasureT22b
2026-10-04 01:37 AEST✓ passGLMTTFTik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2ttft_prefill_tps 201 tok/sT22b
2026-10-04 01:19 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2lab_decode_c1_tps 8.20 tok/s · lab_prefill_tps 147 tok/sgatesT22b
2026-10-04 01:18 AEST✓ passGLMloadik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2load_seconds 47.7 sT22b
2026-10-04 01:01 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2lab_decode_c1_tps 9.50 tok/s · lab_prefill_tps 151 tok/sgatesT22b
2026-10-04 00:51 AEST! failGLMstabilityik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2decode_window_min_tps 9.46 tok/s · tool_calls 0 callsT22b
2026-10-04 00:37 AEST✓ passGLMTTFTik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2ttft_prefill_tps 203 tok/sT22b
2026-10-04 00:22 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2lab_decode_c1_tps 10.0 tok/s · lab_prefill_tps 151 tok/sgatesT22b
2026-10-04 00:22 AEST✓ passGLMloadik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2load_seconds 46.1 sT22b
2026-10-03 23:31 AEST! failGLMlabllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2lab_decode_c1_tps 9.30 tok/s · lab_prefill_tps 80 tok/sgatesT20b
2026-10-03 23:21 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2decode_window_min_tps 8.80 tok/s · tool_calls 19 callsT20b
2026-10-03 22:50 AEST✓ passGLMTTFTllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2ttft_prefill_tps 84.4 tok/sT20b
2026-10-03 21:59 AEST! failGLMlabllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2lab_decode_c1_tps 9.20 tok/s · lab_prefill_tps 78 tok/sgatesT20b
2026-10-03 21:58 AEST✓ passGLMloadllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2load_seconds 49.0 sT20b
2026-10-03 12:48 AEST! failGLMlabglm-llamacpp-iq2m-128k-1x-gpu2lab_decode_c1_tps 9.60 tok/s · lab_prefill_tps 81 tok/sgatesT20b
2026-10-03 12:38 AEST! failGLMstabilityglm-llamacpp-iq2m-128k-1x-gpu2decode_window_min_tps 8.91 tok/s · tool_calls 24 callsT20b
2026-10-03 12:07 AEST✓ passGLMTTFTglm-llamacpp-iq2m-128k-1x-gpu2ttft_prefill_tps 86.3 tok/sT20b
2026-10-03 11:41 AEST! failGLMlabglm-llamacpp-iq2m-128k-1x-gpu2lab_decode_c1_tps 9.60 tok/s · lab_prefill_tps 81 tok/sgatesT20b
2026-10-03 11:41 AEST✓ passGLMloadglm-llamacpp-iq2m-128k-1x-gpu2load_seconds 45.9 sT20b
2026-10-03 11:14 AEST✓ passFlash-Nextlabfn-llamacpp-iq4xs-128k-1x-fit-gpu2lab_decode_c1_tps 26.1 tok/s · lab_prefill_tps 157 tok/sT11
2026-10-03 11:03 AEST! failFlash-Nextstabilityfn-llamacpp-iq4xs-128k-1x-fit-gpu2decode_window_min_tps 23.8 tok/s · tool_calls 49 callsT11
2026-10-03 10:49 AEST✓ passFlash-NextTTFTfn-llamacpp-iq4xs-128k-1x-fit-gpu2ttft_prefill_tps 269 tok/sT11
2026-10-03 10:29 AEST✓ passFlash-Nextlabfn-llamacpp-iq4xs-128k-1x-fit-gpu2lab_decode_c1_tps 26.5 tok/s · lab_prefill_tps 157 tok/sT11
2026-10-03 10:29 AEST✓ passFlash-Nextloadfn-llamacpp-iq4xs-128k-1x-fit-gpu2load_seconds 27.8 sT11
2026-10-03 10:28 AEST✕ crashFlash-Nextloadfn-llamacpp-iq4xs-128k-1x-gpu2load_seconds unavailableloadT11
2026-10-03 08:17 AEST✓ passFlash-Nextlabfn-trellis-3.05-0xsero-main-x8-gpu2lab_decode_c1_tps 55.3 tok/s · lab_prefill_tps 2,228 tok/sT10
2026-10-03 08:06 AEST! failFlash-Nextstabilityfn-trellis-3.05-0xsero-main-x8-gpu2decode_window_min_tps 38.1 tok/s · tool_calls 63 callsT10
2026-10-03 08:06 AEST✓ passFlash-Nextquality panelfn-trellis-3.05-0xsero-main-x8-gpu2top1_agreement 0.99234 fraction · mean_kl 0.0009667 natsT10
2026-10-03 08:04 AEST✓ passFlash-NextTTFTfn-trellis-3.05-0xsero-main-x8-gpu2ttft_prefill_tps 2,263 tok/sT10
2026-10-03 07:48 AEST✓ passFlash-Nextkit sweepfn-trellis-3.05-0xsero-main-x8-gpu2decode_c1_tps 46.9 tok/s · prefill_32k_tps 2,433 tok/sT10
2026-10-03 07:43 AEST✓ passFlash-Nextlabfn-trellis-3.05-0xsero-main-x8-gpu2lab_decode_c1_tps 49.3 tok/s · lab_prefill_tps 2,176 tok/sT10
2026-10-03 07:40 AEST✓ passFlash-Nextloadfn-trellis-3.05-0xsero-main-x8-gpu2load_seconds 179 sT10
2026-10-03 06:54 AEST✓ passnonenvmenvme-rand-1tb-ngramrand_read_iops 168,600 IOPS · rand_read_gbps 3.35 GB/sT05a
2026-10-03 06:53 AEST✓ passnonehw bwhw-bw-gpu0-currenth2d_gbps 6.03 GB/s · d2h_gbps 6.60 GB/sT05a
2026-10-03 06:53 AEST✓ passnonehw bwhw-bw-gpu2-currenth2d_gbps 13.5 GB/s · d2h_gbps 13.2 GB/sT05a

Runs / none

20261002T205342Z-hw-bw-gpu2-current-hw_bw

✓ pass hw bw T05a eligibility: none

Confighw-bw-gpu2-current
Started2026-10-03 06:53 AEST
Finished2026-10-03 06:53 AEST (0.1 min)
Exit status0
Notes1 GiB pinned transfers; full table in detail

Launch

no container: static measurement

Metrics

context.configured–
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
h2d_gbps13.5 GB/s {"1MB_pin": 13.15, "1MB_page": 13.02, "4MB_pin": 13.39, "4MB_page": 13.14, "16MB_pin": 13.45, "16MB_page": 13.19, "64MB_pin": 13.46, "64MB_page": 13.18, "256MB_pin": 13.47, "256MB_page": 13.16, "1024MB_pin": 13.47, "1024MB_page": 13.11}
d2h_gbps13.2 GB/s {"1MB_pin": 12.98, "1MB_page": 7.19, "4MB_pin": 13.14, "4MB_page": 10.06, "16MB_pin": 13.2, "16MB_page": 11.35, "64MB_pin": 13.21, "64MB_page": 12.57, "256MB_pin": 13.21, "256MB_page": 12.61, "1024MB_pin": 13.22, "1024MB_page": 12.69}
host_memcpy_gbps16.6 GB/s

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit483c32c7f11d39f2f1e84c87602db3e166dc9fff clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/none/20261002T205342Z-hw-bw-gpu2-current-hw_bw.json

Runs / none

20261002T205350Z-hw-bw-gpu0-current-hw_bw

✓ pass hw bw T05a eligibility: none

Confighw-bw-gpu0-current
Started2026-10-03 06:53 AEST
Finished2026-10-03 06:54 AEST (0.2 min)
Exit status0
Notes1 GiB pinned transfers; full table in detail

Launch

no container: static measurement

Metrics

context.configured–
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
h2d_gbps6.03 GB/s {"1MB_pin": 6.48, "1MB_page": 6.08, "4MB_pin": 6.6, "4MB_page": 6.1, "16MB_pin": 6.57, "16MB_page": 6.11, "64MB_pin": 6.09, "64MB_page": 6.05, "256MB_pin": 6.03, "256MB_page": 6.01, "1024MB_pin": 6.03, "1024MB_page": 5.99}
d2h_gbps6.60 GB/s {"1MB_pin": 6.51, "1MB_page": 4.62, "4MB_pin": 6.58, "4MB_page": 5.88, "16MB_pin": 6.59, "16MB_page": 6.34, "64MB_pin": 6.6, "64MB_page": 6.48, "256MB_pin": 6.6, "256MB_page": 6.51, "1024MB_pin": 6.6, "1024MB_page": 6.52}
host_memcpy_gbps16.6 GB/s

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit483c32c7f11d39f2f1e84c87602db3e166dc9fff clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/none/20261002T205350Z-hw-bw-gpu0-current-hw_bw.json

Runs / none

20261002T205409Z-nvme-rand-1tb-ngram-nvme_rand

✓ pass nvme T05a eligibility: none

Confignvme-rand-1tb-ngram
Started2026-10-03 06:54 AEST
Finished2026-10-03 06:55 AEST (1.1 min)
Exit status0
Notes~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5/ngram_embedding.safetensors

Launch

no container: static measurement

Metrics

context.configured–
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
rand_read_iops168,600 IOPS "16 KiB, 32 threads"
rand_read_gbps3.35 GB/s [{"bs": 4096, "threads": 1, "GBps": 0.075, "kIOPS": 18.2}, {"bs": 4096, "threads": 8, "GBps": 0.502, "kIOPS": 122.6}, {"bs": 4096, "threads": 32, "GBps": 0.714, "kIOPS": 174.4}, {"bs": 4096, "threads": 64, "GBps": 0.697, "kIOPS": 170.3}, {"bs": 16384, "threads": 1, "GBps": 0.2, "kIOPS": 12.2}, {"bs": 16384, "threads": 8, "GBps": 1.328, "kIOPS": 81.1}, {"bs": 16384, "threads": 32, "GBps": 2.763, "kIOPS": 168.6}, {"bs": 16384, "threads": 64, "GBps": 2.678, "kIOPS": 163.4}, {"bs": 65536, "threads": 1, "GBps": 0.47, "kIOPS": 7.2}, {"bs": 65536, "threads": 8, "GBps": 2.96, "kIOPS": 45.2}, {"bs": 65536, "threads": 32, "GBps": 3.35, "kIOPS": 51.1}, {"bs": 65536, "threads": 64, "GBps": 3.35, "kIOPS": 51.1}, {"bs": 1048576, "threads": 1, "GBps": 1.972, "kIOPS": 1.9}, {"bs": 1048576, "threads": 8, "GBps": 3.35, "kIOPS": 3.2}, {"bs": 1048576, "threads": 32, "GBps": 3.353, "kIOPS": 3.2}, {"bs": 1048576, "threads": 64, "GBps": 3.355, "kIOPS": 3.2}]

Cache, layout and host

Cachewarm
Layoutcurrent · 0 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit483c32c7f11d39f2f1e84c87602db3e166dc9fff clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/none/20261002T205409Z-nvme-rand-1tb-ngram-nvme_rand.json

Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2

20261002T214031Z-fn-trellis-3.05-0xsero-main-x8-gpu2-load

✓ pass load T10 eligibility: none

Configfn-trellis-3.05-0xsero-main-x8-gpu2
Started2026-10-03 07:40 AEST
Finished2026-10-03 07:43 AEST (3.0 min)
Exit statusNone
Notes–

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt)
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Digestsha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds179 s
host_ram_drop_gb63.7 GB
vram_ready_mib21,060 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit966c1355dde3774c7406f9d18e6b8468d4a10fdb clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261002T214031Z-fn-trellis-3.05-0xsero-main-x8-gpu2-load.json

Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2

20261002T214330Z-fn-trellis-3.05-0xsero-main-x8-gpu2-lab

✓ pass lab T10 eligibility: none

Configfn-trellis-3.05-0xsero-main-x8-gpu2
Started2026-10-03 07:43 AEST
Finished2026-10-03 07:48 AEST (5.3 min)
Exit status0
Notesserved=flashnext

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt)
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Digestsha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max166,667 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps49.3 tok/s
lab_prefill_tps2,176 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit966c1355dde3774c7406f9d18e6b8468d4a10fdb clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261002T214330Z-fn-trellis-3.05-0xsero-main-x8-gpu2-lab.json

Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2

20261002T214848Z-fn-trellis-3.05-0xsero-main-x8-gpu2-kit_sweep

✓ pass kit sweep T10 eligibility: none

Configfn-trellis-3.05-0xsero-main-x8-gpu2
Started2026-10-03 07:48 AEST
Finished2026-10-03 08:04 AEST (15.9 min)
Exit status0
Notessweep status: DONE

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt)
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Digestsha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
prefill_8k_tps2,292 tok/s {"min": 2277.9, "max": 2305.2, "n": 3}
prefill_16k_tps2,401 tok/s {"min": 2400.6, "max": 2403.1, "n": 3}
prefill_32k_tps2,433 tok/s {"min": 2431.6, "max": 2433.5, "n": 3}
prefill_64k_tps2,415 tok/s {"min": 2413.8, "max": 2415.0, "n": 3}
decode_c1_tps46.9 tok/s {"min": 45.1, "max": 48.7, "power_w": null}
decode_c2_tps51.6 tok/s {"min": 51.1, "max": 52.11, "power_w": null}
decode_c3_tps53.9 tok/s {"min": 51.82, "max": 55.98, "power_w": null}
decode_c4_tps51.8 tok/s {"min": 50.57, "max": 53.12, "power_w": null}
decode_c1_32k_tps48.5 tok/s {"min": 48.51, "max": 48.6, "power_w": null}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit966c1355dde3774c7406f9d18e6b8468d4a10fdb clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261002T214848Z-fn-trellis-3.05-0xsero-main-x8-gpu2-kit_sweep.json

Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2

20261002T220441Z-fn-trellis-3.05-0xsero-main-x8-gpu2-ttft

✓ pass TTFT T10 eligibility: none

Configfn-trellis-3.05-0xsero-main-x8-gpu2
Started2026-10-03 08:04 AEST
Finished2026-10-03 08:06 AEST (1.5 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt)
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Digestsha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps2,263 tok/s {"prompt_tokens": [13679, 13707, 13790], "ttft_s": [6.049, 6.056, 6.045]}
ttft_prefill_32k_tps2,368 tok/s {"prompt_tokens": [54780, 54797, 54620], "ttft_s": [23.132, 23.12, 23.139]}
ttft_prefill_tps2,263 tok/s {"prompt_tokens": [13679, 13707, 13790], "ttft_s": [6.049, 6.056, 6.045]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit966c1355dde3774c7406f9d18e6b8468d4a10fdb clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261002T220441Z-fn-trellis-3.05-0xsero-main-x8-gpu2-ttft.json

Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2

20261002T220609Z-fn-trellis-3.05-0xsero-main-x8-gpu2-quality_panel

✓ pass quality panel T10 eligibility: none

Configfn-trellis-3.05-0xsero-main-x8-gpu2
Started2026-10-03 08:06 AEST
Finished2026-10-03 08:06 AEST (0.4 min)
Exit status0
Notespanel=qwen3.8-flash-next-exl3-ref-panel.json rc=0

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt)
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Digestsha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
top1_agreement0.99234 fraction
mean_kl0.0009667 nats
in_bandyes "top1>=0.987, KL<=0.0013"

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit966c1355dde3774c7406f9d18e6b8468d4a10fdb clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261002T220609Z-fn-trellis-3.05-0xsero-main-x8-gpu2-quality_panel.json

Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2

20261002T220633Z-fn-trellis-3.05-0xsero-main-x8-gpu2-stability

! fail stability T10 eligibility: none

Configfn-trellis-3.05-0xsero-main-x8-gpu2
Started2026-10-03 08:06 AEST
Finished2026-10-03 08:17 AEST (10.5 min)
Exit status0
Notesstability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 63 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 38.077 < floor 50.0

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt)
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Digestsha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max30,034 tokens
context.headroom_min174,233 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls63 calls {"per_min": 6.3, "requests": 45}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /v1/tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps38.1 tok/s
window 1: 42.97 tok/swindow 2: 43.5 tok/swindow 3: 47.28 tok/swindow 4: 41.43 tok/swindow 5: 43.59 tok/swindow 6: 42.67 tok/smin 41.43 · max 47.28 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 331, "min_at_generation_s": 63.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps43.1 tok/s {"generation_seconds": 450.481, "generated_tokens": 19411.4}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 208.0, "requests": 18, "tool_calls": 26, "completion_tokens": 6387, "occupied_max": 12502}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 392.0, "requests": 27, "tool_calls": 37, "completion_tokens": 13133, "occupied_max": 30034}] list
tasks_over_64k[] instance ids
completion_tokens19,520 tokens {"finish_reasons": {"tool_calls": 45}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 63 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 38.1 tok/s, whole run 43.1 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit966c1355dde3774c7406f9d18e6b8468d4a10fdb clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261002T220633Z-fn-trellis-3.05-0xsero-main-x8-gpu2-stability.json

Runs / Qwen3.8-Flash-Next / fn-trellis-3.05-0xsero-main-x8-gpu2

20261002T221703Z-fn-trellis-3.05-0xsero-main-x8-gpu2-lab

✓ pass lab T10 eligibility: none

Configfn-trellis-3.05-0xsero-main-x8-gpu2
Started2026-10-03 08:17 AEST
Finished2026-10-03 08:22 AEST (5.0 min)
Exit status0
Notesserved=flashnext

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261002T214031Z-lab-fn-trellis-3.05-0xsero-main-x8-gpu2/docker-run.txt)
docker run -d --name lab-fn-trellis-3.05-0xsero-main-x8-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Digestsha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.json sha256 248d18e53c5004810964b493e2676bb465a9386a5c527a8b88e67cbfd0dffa68
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max166,667 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps55.3 tok/s
lab_prefill_tps2,228 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit966c1355dde3774c7406f9d18e6b8468d4a10fdb clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261002T221703Z-fn-trellis-3.05-0xsero-main-x8-gpu2-lab.json

Runs / Qwen3.8-Flash-Next / fn-llamacpp-iq4xs-128k-1x-gpu2

20261003T002815Z-fn-llamacpp-iq4xs-128k-1x-gpu2-load

✕ crash load T11 eligibility: none failed stage: load

Configfn-llamacpp-iq4xs-128k-1x-gpu2
Started2026-10-03 10:28 AEST
Finished2026-10-03 10:28 AEST (0.4 min)
Exit status1
Notesuser, abort 0.18.994.031 E ggml_backend_cuda_buffer_type_alloc_buffer: allocating 65358.17 MiB on device 0: cudaMalloc failed: out of memory 0.18.994.037 E alloc_tensor_range: failed to allocate CUDA0 buffer of size 68533006336 0.19.139.695 E llama_model_load: error loading model: unable to allocate CUDA0 buffer 0.19.139.702 E llama_model_load_from_file_impl: failed to load model 0.19.139.708 E cmn common_init_: failed to load model '/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf' 0.19.139.713 E srv load_model: failed to load model, '/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf' 0.19.139.715 I srv operator(): operator(): cleaning up before exit... 0.19.140.570 E srv llama_server: exiting due to model loading error

Launch

recorded by tools/run.py (artifacts/runs/20261003T002815Z-lab-fn-llamacpp-iq4xs-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-fn-llamacpp-iq4xs-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 -ot per_layer_token_embd.weight=CPU
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-1x.json sha256 697275c75292e4ef45a02a99b2a98c23b7a55630ed0430af0972a1dcf84b47d8
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 -ot per_layer_token_embd.weight=CPU
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_secondsunavailable not produced (crash/load)
host_ram_drop_gbunavailable not produced (crash/load)
vram_ready_mibunavailable not produced (crash/load)

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit1fa796f9d509705f24dc1870357869cf7c293bc2 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261003T002815Z-fn-llamacpp-iq4xs-128k-1x-gpu2-load.json

Runs / Qwen3.8-Flash-Next / fn-llamacpp-iq4xs-128k-1x-fit-gpu2

20261003T002901Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-load

✓ pass load T11 eligibility: none

Configfn-llamacpp-iq4xs-128k-1x-fit-gpu2
Started2026-10-03 10:29 AEST
Finished2026-10-03 10:29 AEST (0.5 min)
Exit statusNone
Notes–

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T002901Z-lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2/docker-run.txt)
docker run -d --name lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-1x-fit.json sha256 67f1655622a5b08eae6fb8c89ca4a189aab6f7578a1a740baae4832ddc4e8160
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds27.8 s
host_ram_drop_gb2.30 GB
vram_ready_mib22,684 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit1fa796f9d509705f24dc1870357869cf7c293bc2 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261003T002901Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-load.json

Runs / Qwen3.8-Flash-Next / fn-llamacpp-iq4xs-128k-1x-fit-gpu2

20261003T002929Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-lab

✓ pass lab T11 eligibility: none

Configfn-llamacpp-iq4xs-128k-1x-fit-gpu2
Started2026-10-03 10:29 AEST
Finished2026-10-03 10:49 AEST (19.9 min)
Exit status0
Notesserved=/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T002901Z-lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2/docker-run.txt)
docker run -d --name lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-1x-fit.json sha256 67f1655622a5b08eae6fb8c89ca4a189aab6f7578a1a740baae4832ddc4e8160
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_max106,294 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps26.5 tok/s
lab_prefill_tps157 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit1fa796f9d509705f24dc1870357869cf7c293bc2 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261003T002929Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-lab.json

Runs / Qwen3.8-Flash-Next / fn-llamacpp-iq4xs-128k-1x-fit-gpu2

20261003T004923Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-ttft

✓ pass TTFT T11 eligibility: none

Configfn-llamacpp-iq4xs-128k-1x-fit-gpu2
Started2026-10-03 10:49 AEST
Finished2026-10-03 11:03 AEST (14.5 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T002901Z-lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2/docker-run.txt)
docker run -d --name lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-1x-fit.json sha256 67f1655622a5b08eae6fb8c89ca4a189aab6f7578a1a740baae4832ddc4e8160
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps269 tok/s {"prompt_tokens": [13827, 13626, 13781], "ttft_s": [50.985, 51.655, 51.198]}
ttft_prefill_32k_tps230 tok/s {"prompt_tokens": [54692, 54711, 54604], "ttft_s": [237.983, 237.379, 240.167]}
ttft_prefill_tps269 tok/s {"prompt_tokens": [13827, 13626, 13781], "ttft_s": [50.985, 51.655, 51.198]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit1fa796f9d509705f24dc1870357869cf7c293bc2 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261003T004923Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-ttft.json

Runs / Qwen3.8-Flash-Next / fn-llamacpp-iq4xs-128k-1x-fit-gpu2

20261003T010352Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-stability

! fail stability T11 eligibility: none

Configfn-llamacpp-iq4xs-128k-1x-fit-gpu2
Started2026-10-03 11:03 AEST
Finished2026-10-03 11:14 AEST (10.5 min)
Exit status0
Notesstability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 49 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 1 repetition hit(s); decode window min 23.83 < floor 50.0

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T002901Z-lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2/docker-run.txt)
docker run -d --name lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-1x-fit.json sha256 67f1655622a5b08eae6fb8c89ca4a189aab6f7578a1a740baae4832ddc4e8160
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_max16,819 tokens
context.headroom_min114,087 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls49 calls {"per_min": 4.9, "requests": 32}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits1 hits {"ngram64x3": 1, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps23.8 tok/s
window 1: 24.81 tok/swindow 2: 24.06 tok/swindow 3: 25.42 tok/swindow 4: 24.9 tok/swindow 5: 24.6 tok/smin 24.06 · max 25.42 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 271, "min_at_generation_s": 105.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps24.8 tok/s {"generation_seconds": 390.845, "generated_tokens": 9687.8}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 306.9, "requests": 17, "tool_calls": 26, "completion_tokens": 4940, "occupied_max": 16819}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 293.1, "requests": 16, "tool_calls": 24, "completion_tokens": 4788, "occupied_max": 12213}] list
tasks_over_64k[] instance ids
completion_tokens9,728 tokens {"finish_reasons": {"tool_calls": 32}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 49 tool calls, 0 parse failures, 1 repetition hits, tasks solved 1/2, decode min window 23.8 tok/s, whole run 24.8 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit1fa796f9d509705f24dc1870357869cf7c293bc2 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261003T010352Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-stability.json

Runs / Qwen3.8-Flash-Next / fn-llamacpp-iq4xs-128k-1x-fit-gpu2

20261003T011420Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-lab

✓ pass lab T11 eligibility: none

Configfn-llamacpp-iq4xs-128k-1x-fit-gpu2
Started2026-10-03 11:14 AEST
Finished2026-10-03 11:40 AEST (26.5 min)
Exit status0
Notesserved=/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T002901Z-lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2/docker-run.txt)
docker run -d --name lab-fn-llamacpp-iq4xs-128k-1x-fit-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-1x-fit.json sha256 67f1655622a5b08eae6fb8c89ca4a189aab6f7578a1a740baae4832ddc4e8160
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_max106,294 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps26.1 tok/s
lab_prefill_tps157 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit1fa796f9d509705f24dc1870357869cf7c293bc2 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261003T011420Z-fn-llamacpp-iq4xs-128k-1x-fit-gpu2-lab.json

Runs / GLM-5.3-Flash / glm-llamacpp-iq2m-128k-1x-gpu2

20261003T014103Z-glm-llamacpp-iq2m-128k-1x-gpu2-load

✓ pass load T20b eligibility: none

Configglm-llamacpp-iq2m-128k-1x-gpu2
Started2026-10-03 11:41 AEST
Finished2026-10-03 11:41 AEST (0.8 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T014103Z-lab-glm-llamacpp-iq2m-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-glm-llamacpp-iq2m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x.json sha256 01715887da6d653bd4cc63d959d62744ec8f44f30ca1a60627052180981e3fc1
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds45.9 s
host_ram_drop_gb2.23 GB
vram_ready_mib22,784 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit4b12eecb28cf351fc09677fc509debe454394308 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T014103Z-glm-llamacpp-iq2m-128k-1x-gpu2-load.json

Runs / GLM-5.3-Flash / glm-llamacpp-iq2m-128k-1x-gpu2

20261003T014149Z-glm-llamacpp-iq2m-128k-1x-gpu2-lab

! fail lab T20b eligibility: none failed stage: gates

Configglm-llamacpp-iq2m-128k-1x-gpu2
Started2026-10-03 11:41 AEST
Finished2026-10-03 12:07 AEST (26.0 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T014103Z-lab-glm-llamacpp-iq2m-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-glm-llamacpp-iq2m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x.json sha256 01715887da6d653bd4cc63d959d62744ec8f44f30ca1a60627052180981e3fc1
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps9.60 tok/s
lab_prefill_tps81 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit4b12eecb28cf351fc09677fc509debe454394308 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T014149Z-glm-llamacpp-iq2m-128k-1x-gpu2-lab.json

Runs / GLM-5.3-Flash / glm-llamacpp-iq2m-128k-1x-gpu2

20261003T020751Z-glm-llamacpp-iq2m-128k-1x-gpu2-ttft

✓ pass TTFT T20b eligibility: none

Configglm-llamacpp-iq2m-128k-1x-gpu2
Started2026-10-03 12:07 AEST
Finished2026-10-03 12:38 AEST (30.5 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T014103Z-lab-glm-llamacpp-iq2m-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-glm-llamacpp-iq2m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x.json sha256 01715887da6d653bd4cc63d959d62744ec8f44f30ca1a60627052180981e3fc1
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps86.3 tok/s {"prompt_tokens": [10742, 10816, 10721], "ttft_s": [124.451, 126.526, 123.598]}
ttft_prefill_32k_tps88.5 tok/s {"prompt_tokens": [42736, 42800, 42830], "ttft_s": [485.781, 483.408, 483.895]}
ttft_prefill_tps86.3 tok/s {"prompt_tokens": [10742, 10816, 10721], "ttft_s": [124.451, 126.526, 123.598]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbff973e82a9b895daba523707af9c56454e76a39 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T020751Z-glm-llamacpp-iq2m-128k-1x-gpu2-ttft.json

Runs / GLM-5.3-Flash / glm-llamacpp-iq2m-128k-1x-gpu2

20261003T023819Z-glm-llamacpp-iq2m-128k-1x-gpu2-stability

! fail stability T20b eligibility: none

Configglm-llamacpp-iq2m-128k-1x-gpu2
Started2026-10-03 12:38 AEST
Finished2026-10-03 12:48 AEST (10.1 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 24 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 8.91 < floor 15.0

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T014103Z-lab-glm-llamacpp-iq2m-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-glm-llamacpp-iq2m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x.json sha256 01715887da6d653bd4cc63d959d62744ec8f44f30ca1a60627052180981e3fc1
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max9,047 tokens
context.headroom_min121,785 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls24 calls {"per_min": 2.4, "requests": 13}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps8.91 tok/s
window 1: 8.98 tok/swindow 2: 9.23 tok/swindow 3: 9.09 tok/swindow 4: 9.18 tok/swindow 5: 9.04 tok/swindow 6: 9.14 tok/smin 8.98 · max 9.23 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 328, "min_at_generation_s": 90.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps9.14 tok/s {"generation_seconds": 447.593, "generated_tokens": 4091.5}
tasks_attempted1 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved0 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 600.0, "requests": 14, "tool_calls": 24, "completion_tokens": 4105, "occupied_max": 9047}] list
tasks_over_64k[] instance ids
completion_tokens4,105 tokens {"finish_reasons": {"tool_calls": 13}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 24 tool calls, 0 parse failures, 0 repetition hits, tasks solved 0/1, decode min window 8.91 tok/s, whole run 9.14 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbff973e82a9b895daba523707af9c56454e76a39 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T023819Z-glm-llamacpp-iq2m-128k-1x-gpu2-stability.json

Runs / GLM-5.3-Flash / glm-llamacpp-iq2m-128k-1x-gpu2

20261003T024824Z-glm-llamacpp-iq2m-128k-1x-gpu2-lab

! fail lab T20b eligibility: none failed stage: gates

Configglm-llamacpp-iq2m-128k-1x-gpu2
Started2026-10-03 12:48 AEST
Finished2026-10-03 13:14 AEST (25.9 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T014103Z-lab-glm-llamacpp-iq2m-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-glm-llamacpp-iq2m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x.json sha256 01715887da6d653bd4cc63d959d62744ec8f44f30ca1a60627052180981e3fc1
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps9.60 tok/s
lab_prefill_tps81 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbff973e82a9b895daba523707af9c56454e76a39 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T024824Z-glm-llamacpp-iq2m-128k-1x-gpu2-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2

20261003T115813Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-load

✓ pass load T20b eligibility: none

Configllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2
Started2026-10-03 21:58 AEST
Finished2026-10-03 21:59 AEST (0.8 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T115813Z-lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-Q2_K:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-q2-k-128k-1x.json sha256 f7ab285db47877ab680a678ead1d3de7960b1a54121d89c804a51dc6c2477de0
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-Q2_K.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds49.0 s
host_ram_drop_gb2.22 GB
vram_ready_mib22,626 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit73bf82fd3254285cf46014dbef7e02e6ef6db9b3 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T115813Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2

20261003T115902Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-lab

! fail lab T20b eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2
Started2026-10-03 21:59 AEST
Finished2026-10-03 22:50 AEST (51.0 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T115813Z-lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-Q2_K:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-q2-k-128k-1x.json sha256 f7ab285db47877ab680a678ead1d3de7960b1a54121d89c804a51dc6c2477de0
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-Q2_K.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps9.20 tok/s
lab_prefill_tps78 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbc100ca481d268577db82a7adcd92133db2241e4 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T115902Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2

20261003T125002Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-ttft

✓ pass TTFT T20b eligibility: none

Configllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2
Started2026-10-03 22:50 AEST
Finished2026-10-03 23:21 AEST (31.5 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T115813Z-lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-Q2_K:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-q2-k-128k-1x.json sha256 f7ab285db47877ab680a678ead1d3de7960b1a54121d89c804a51dc6c2477de0
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-Q2_K.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps84.4 tok/s {"prompt_tokens": [10644, 10721, 10688], "ttft_s": [126.044, 127.442, 125.735]}
ttft_prefill_32k_tps84.9 tok/s {"prompt_tokens": [42901, 42759, 42872], "ttft_s": [504.114, 504.375, 505.138]}
ttft_prefill_tps84.4 tok/s {"prompt_tokens": [10644, 10721, 10688], "ttft_s": [126.044, 127.442, 125.735]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbc100ca481d268577db82a7adcd92133db2241e4 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T125002Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-ttft.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2

20261003T132135Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-stability

! fail stability T20b eligibility: none

Configllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2
Started2026-10-03 23:21 AEST
Finished2026-10-03 23:31 AEST (10.1 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 19 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 8.8 < floor 15.0

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T115813Z-lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-Q2_K:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-q2-k-128k-1x.json sha256 f7ab285db47877ab680a678ead1d3de7960b1a54121d89c804a51dc6c2477de0
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-Q2_K.json

Metrics

context.configured131,072 tokens
context.occupied_max6,658 tokens
context.headroom_min123,658 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls19 calls {"per_min": 1.9, "requests": 11}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps8.80 tok/s
window 1: 8.89 tok/swindow 2: 8.87 tok/swindow 3: 8.93 tok/swindow 4: 8.88 tok/swindow 5: 9.0 tok/smin 8.87 · max 9 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 267, "min_at_generation_s": 131.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps8.93 tok/s {"generation_seconds": 386.503, "generated_tokens": 3451.5}
tasks_attempted1 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved0 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 600.0, "requests": 12, "tool_calls": 20, "completion_tokens": 3463, "occupied_max": 6658}] list
tasks_over_64k[] instance ids
completion_tokens3,463 tokens {"finish_reasons": {"tool_calls": 11}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 19 tool calls, 0 parse failures, 0 repetition hits, tasks solved 0/1, decode min window 8.80 tok/s, whole run 8.93 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbc100ca481d268577db82a7adcd92133db2241e4 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T132135Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-stability.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2

20261003T133139Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-lab

! fail lab T20b eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2
Started2026-10-03 23:31 AEST
Finished2026-10-04 00:21 AEST (50.3 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T115813Z-lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-Q2_K:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-q2-k-128k-1x.json sha256 f7ab285db47877ab680a678ead1d3de7960b1a54121d89c804a51dc6c2477de0
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-Q2_K.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps9.30 tok/s
lab_prefill_tps80 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbc100ca481d268577db82a7adcd92133db2241e4 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T133139Z-llama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2-lab.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2

20261003T142203Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-load

✓ pass load T22b eligibility: none

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2
Started2026-10-04 00:22 AEST
Finished2026-10-04 00:22 AEST (0.8 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T142203Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x.json sha256 168e8b490c871dcebe6afed1dfd973be9985379eb748e37f17be98b86a86e44e
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds46.1 s
host_ram_drop_gb111 GB
vram_ready_mib14,936 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit57320e8fbf776bca99fc63e2de1647d22ef83f06 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T142203Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-load.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2

20261003T142250Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-lab

! fail lab T22b eligibility: none failed stage: gates

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2
Started2026-10-04 00:22 AEST
Finished2026-10-04 00:37 AEST (14.7 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T142203Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x.json sha256 168e8b490c871dcebe6afed1dfd973be9985379eb748e37f17be98b86a86e44e
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}
lab_decode_c1_tps10.0 tok/s
lab_prefill_tps151 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit57320e8fbf776bca99fc63e2de1647d22ef83f06 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T142250Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-lab.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2

20261003T143733Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-ttft

✓ pass TTFT T22b eligibility: none

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2
Started2026-10-04 00:37 AEST
Finished2026-10-04 00:51 AEST (13.5 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T142203Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x.json sha256 168e8b490c871dcebe6afed1dfd973be9985379eb748e37f17be98b86a86e44e
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps203 tok/s {"prompt_tokens": [10751, 10785, 10716], "ttft_s": [52.615, 53.39, 52.749]}
ttft_prefill_32k_tps197 tok/s {"prompt_tokens": [42739, 42927, 42933], "ttft_s": [217.584, 218.259, 218.191]}
ttft_prefill_tps203 tok/s {"prompt_tokens": [10751, 10785, 10716], "ttft_s": [52.615, 53.39, 52.749]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit57320e8fbf776bca99fc63e2de1647d22ef83f06 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T143733Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-ttft.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2

20261003T145106Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-stability

! fail stability T22b eligibility: none

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2
Started2026-10-04 00:51 AEST
Finished2026-10-04 01:01 AEST (10.1 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 47 tool-call parse failure(s); decode window min 9.458 < floor 15.0

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T142203Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x.json sha256 168e8b490c871dcebe6afed1dfd973be9985379eb748e37f17be98b86a86e44e
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max1,874 tokens
context.headroom_min129,125 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls0 calls {"per_min": 0.0, "requests": 47}
parse_failures47 responses {"wire": {"unparsed_markup": 47}, "harness_format_errors": {"no_tool_call": 47}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors47 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps9.46 tok/s
window 1: 9.49 tok/swindow 2: 9.55 tok/swindow 3: 9.5 tok/swindow 4: 9.46 tok/swindow 5: 9.55 tok/smin 9.46 · max 9.55 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 277, "min_at_generation_s": 144.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps9.50 tok/s {"generation_seconds": 396.975, "generated_tokens": 3772.8}
tasks_attempted16 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved0 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 51.1, "requests": 3, "tool_calls": 0, "completion_tokens": 310, "occupied_max": 1559}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 40.8, "requests": 3, "tool_calls": 0, "completion_tokens": 227, "occupied_max": 1641}, {"attempt": 3, "round": 1, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 39.7, "requests": 3, "tool_calls": 0, "completion_tokens": 220, "occupied_max": 1609}, {"attempt": 4, "round": 1, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 43.5, "requests": 3, "tool_calls": 0, "completion_tokens": 251, "occupied_max": 1874}, {"attempt": 5, "round": 1, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 41.3, "requests": 3, "tool_calls": 0, "completion_tokens": 233, "occupied_max": 1770}, {"attempt": 6, "round": 1, "instance_id": "astropy__astropy-12907", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 34.9, "requests": 3, "tool_calls": 0, "completion_tokens": 174, "occupied_max": 1716}, {"attempt": 7, "round": 1, "instance_id": "pytest-dev__pytest-7373", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 41.1, "requests": 3, "tool_calls": 0, "completion_tokens": 231, "occupied_max": 1641}, {"attempt": 8, "round": 1, "instance_id": "sympy__sympy-20590", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 52.5, "requests": 3, "tool_calls": 0, "completion_tokens": 339, "occupied_max": 1578}, {"attempt": 9, "round": 2, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 31.5, "requests": 3, "tool_calls": 0, "completion_tokens": 230, "occupied_max": 1559}, {"attempt": 10, "round": 2, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 35.7, "requests": 3, "tool_calls": 0, "completion_tokens": 270, "occupied_max": 1641}, {"attempt": 11, "round": 2, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 32.8, "requests": 3, "tool_calls": 0, "completion_tokens": 243, "occupied_max": 1609}, {"attempt": 12, "round": 2, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 33.5, "requests": 3, "tool_calls": 0, "completion_tokens": 249, "occupied_max": 1874}, {"attempt": 13, "round": 2, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 33.7, "requests": 3, "tool_calls": 0, "completion_tokens": 251, "occupied_max": 1770}, {"attempt": 14, "round": 2, "instance_id": "astropy__astropy-12907", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 25.9, "requests": 3, "tool_calls": 0, "completion_tokens": 178, "occupied_max": 1716}, {"attempt": 15, "round": 2, "instance_id": "pytest-dev__pytest-7373", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 30.0, "requests": 3, "tool_calls": 0, "completion_tokens": 217, "occupied_max": 1641}, {"attempt": 16, "round": 2, "instance_id": "sympy__sympy-20590", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 32.0, "requests": 3, "tool_calls": 0, "completion_tokens": 198, "occupied_max": 1459}] list
tasks_over_64k[] instance ids
completion_tokens3,821 tokens {"finish_reasons": {"stop": 47}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 0 tool calls, 47 parse failures, 0 repetition hits, tasks solved 0/16, decode min window 9.46 tok/s, whole run 9.50 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit57320e8fbf776bca99fc63e2de1647d22ef83f06 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T145106Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-stability.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2

20261003T150111Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-lab

! fail lab T22b eligibility: none failed stage: gates

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2
Started2026-10-04 01:01 AEST
Finished2026-10-04 01:18 AEST (17.7 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T142203Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x.json sha256 168e8b490c871dcebe6afed1dfd973be9985379eb748e37f17be98b86a86e44e
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}
lab_decode_c1_tps9.50 tok/s
lab_prefill_tps151 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit57320e8fbf776bca99fc63e2de1647d22ef83f06 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T150111Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2-lab.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2

20261003T151858Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-load

✓ pass load T22b eligibility: none

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2
Started2026-10-04 01:18 AEST
Finished2026-10-04 01:19 AEST (0.8 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! TTFT stability ✕

Launch

recorded by tools/run.py (artifacts/runs/20261003T151858Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp.json sha256 c2df636d0f6ab990abe7dd3c658b064f3c8df59d819b7fe907ce82a12b84a9c4
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds47.7 s
host_ram_drop_gb115 GB
vram_ready_mib22,318 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit664783f004ff58b6fcabb0691d6b1c8298e2e818 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T151858Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-load.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2

20261003T151947Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-lab

! fail lab T22b eligibility: none failed stage: gates

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2
Started2026-10-04 01:19 AEST
Finished2026-10-04 01:37 AEST (17.7 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! TTFT stability ✕

Launch

recorded by tools/run.py (artifacts/runs/20261003T151858Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp.json sha256 c2df636d0f6ab990abe7dd3c658b064f3c8df59d819b7fe907ce82a12b84a9c4
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}
lab_decode_c1_tps8.20 tok/s
lab_prefill_tps147 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsunavailable not produced (fail/gates)
mtp_acceptance_rateunavailable not produced (fail/gates)

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit937eb11054a0e7abd1b0e970f25ea2fda3f4cff9 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T151947Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-lab.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2

20261003T153729Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-ttft

✓ pass TTFT T22b eligibility: none

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2
Started2026-10-04 01:37 AEST
Finished2026-10-04 01:51 AEST (13.9 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! TTFT stability ✕

Launch

recorded by tools/run.py (artifacts/runs/20261003T151858Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp.json sha256 c2df636d0f6ab990abe7dd3c658b064f3c8df59d819b7fe907ce82a12b84a9c4
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps201 tok/s {"prompt_tokens": [10740, 10677, 10728], "ttft_s": [53.545, 52.784, 53.644]}
ttft_prefill_32k_tps190 tok/s {"prompt_tokens": [42815, 42706, 42534], "ttft_s": [223.459, 224.655, 228.199]}
ttft_prefill_tps201 tok/s {"prompt_tokens": [10740, 10677, 10728], "ttft_s": [53.545, 52.784, 53.644]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit937eb11054a0e7abd1b0e970f25ea2fda3f4cff9 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T153729Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-ttft.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2

20261003T155126Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-stability

✕ crash stability T22b eligibility: none failed stage: measure

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2
Started2026-10-04 01:51 AEST
Finished2026-10-04 02:01 AEST (10.1 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); 9 tool-call parse failure(s); decode floor not verifiable (too little generation)

Other runs of this config: load lab ! TTFT stability ✕

Launch

recorded by tools/run.py (artifacts/runs/20261003T151858Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp.json sha256 c2df636d0f6ab990abe7dd3c658b064f3c8df59d819b7fe907ce82a12b84a9c4
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max1,641 tokens
context.headroom_min129,356 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls0 calls {"per_min": 0.0, "requests": 9}
parse_failures9 responses {"wire": {"unparsed_markup": 9}, "harness_format_errors": {"no_tool_call": 9}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors9 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tpsunavailable only 57.72 s of generation; need > 120 s for one window
decode_whole_run_tps12.6 tok/s {"generation_seconds": 57.72, "generated_tokens": 729.8}
tasks_attempted5 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved0 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomyes {"upstream_connection_failures": 54, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 47.1, "requests": 3, "tool_calls": 0, "completion_tokens": 274, "occupied_max": 1559}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 36.4, "requests": 3, "tool_calls": 0, "completion_tokens": 248, "occupied_max": 1641}, {"attempt": 3, "round": 1, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 33.8, "requests": 3, "tool_calls": 0, "completion_tokens": 217, "occupied_max": 1609}, {"attempt": 4, "round": 1, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "Error:BadGatewayError: litellm.BadGatewayError: BadGatewayError: OpenAIException - Cannot connect to host 127.0.0.1:18080 ssl:default [Connect call failed ('127.0.0.1', 18080)]", "submitted": false, "resolved": false, "seconds": 343.5}, {"attempt": 5, "round": 1, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 139.3}] list
tasks_over_64k[] instance ids
completion_tokens739 tokens {"finish_reasons": {"stop": 9}}
mtp_accepted_tpsunavailable not produced (crash/measure)
mtp_acceptance_rateunavailable not produced (crash/measure)

Agentic summary: 600 s, 0 tool calls, 9 parse failures, 0 repetition hits, tasks solved 0/5, decode min window – tok/s, whole run 12.6 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit937eb11054a0e7abd1b0e970f25ea2fda3f4cff9 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T155126Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2-stability.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2

20261003T162532Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-load

✓ pass load T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2
Started2026-10-04 02:25 AEST
Finished2026-10-04 02:26 AEST (1.0 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T162532Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x.json sha256 76879bccb4223161305f6a73f10b3a4f79db3609c18c993495ef170aa8ae9db8
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds58.2 s
host_ram_drop_gb1.89 GB
vram_ready_mib22,596 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9fc82ec3597413063d86252387463dee533769ed clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T162532Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2

20261003T162631Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-lab

! fail lab T20b eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2
Started2026-10-04 02:26 AEST
Finished2026-10-04 02:55 AEST (28.7 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T162532Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x.json sha256 76879bccb4223161305f6a73f10b3a4f79db3609c18c993495ef170aa8ae9db8
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps8.30 tok/s
lab_prefill_tps65 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9fc82ec3597413063d86252387463dee533769ed clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T162631Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2

20261003T165514Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-ttft

✓ pass TTFT T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2
Started2026-10-04 02:55 AEST
Finished2026-10-04 03:30 AEST (34.8 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T162532Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x.json sha256 76879bccb4223161305f6a73f10b3a4f79db3609c18c993495ef170aa8ae9db8
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps75.9 tok/s {"prompt_tokens": [10736, 10754, 10667], "ttft_s": [140.833, 141.619, 141.761]}
ttft_prefill_32k_tps77.3 tok/s {"prompt_tokens": [42840, 42803, 42831], "ttft_s": [554.48, 553.698, 555.603]}
ttft_prefill_tps75.9 tok/s {"prompt_tokens": [10736, 10754, 10667], "ttft_s": [140.833, 141.619, 141.761]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9fc82ec3597413063d86252387463dee533769ed clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T165514Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-ttft.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2

20261003T173003Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-stability

! fail stability T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2
Started2026-10-04 03:30 AEST
Finished2026-10-04 03:40 AEST (10.5 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 25 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 7.646 < floor 15.0

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T162532Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x.json sha256 76879bccb4223161305f6a73f10b3a4f79db3609c18c993495ef170aa8ae9db8
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json

Metrics

context.configured131,072 tokens
context.occupied_max4,875 tokens
context.headroom_min126,123 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls25 calls {"per_min": 2.5, "requests": 15}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps7.65 tok/s
window 1: 8.08 tok/swindow 2: 7.78 tok/swindow 3: 7.99 tok/swindow 4: 8.16 tok/swindow 5: 8.19 tok/smin 7.78 · max 8.19 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 269, "min_at_generation_s": 156.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps8.08 tok/s {"generation_seconds": 388.925, "generated_tokens": 3143.9}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 377.0, "requests": 10, "tool_calls": 17, "completion_tokens": 2133, "occupied_max": 4875}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 223.0, "requests": 6, "tool_calls": 9, "completion_tokens": 1027, "occupied_max": 4558}] list
tasks_over_64k[] instance ids
completion_tokens3,160 tokens {"finish_reasons": {"tool_calls": 15}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 25 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 7.65 tok/s, whole run 8.08 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9fc82ec3597413063d86252387463dee533769ed clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T173003Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-stability.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2

20261003T174033Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-lab

! fail lab T20b eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2
Started2026-10-04 03:40 AEST
Finished2026-10-04 04:25 AEST (45.4 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T162532Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x.json sha256 76879bccb4223161305f6a73f10b3a4f79db3609c18c993495ef170aa8ae9db8
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps8.40 tok/s
lab_prefill_tps71 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9fc82ec3597413063d86252387463dee533769ed clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T174033Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2

20261003T182602Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-load

✓ pass load T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2
Started2026-10-04 04:26 AEST
Finished2026-10-04 04:26 AEST (0.8 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T182602Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048.json sha256 89ee35a7fb8e6ec8fd63aa7e9e07a5ed1c3448eed3d10cebf357887e810a94fe
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds45.9 s
host_ram_drop_gb3.54 GB
vram_ready_mib21,944 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitd4f093c724ebfd683feed823be35a11d4c009375 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T182602Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2

20261003T182648Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-lab

! fail lab T20b eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2
Started2026-10-04 04:26 AEST
Finished2026-10-04 04:54 AEST (27.5 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T182602Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048.json sha256 89ee35a7fb8e6ec8fd63aa7e9e07a5ed1c3448eed3d10cebf357887e810a94fe
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps9.50 tok/s
lab_prefill_tps200 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitd4f093c724ebfd683feed823be35a11d4c009375 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T182648Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2

20261003T185415Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-ttft

✓ pass TTFT T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2
Started2026-10-04 04:54 AEST
Finished2026-10-04 05:06 AEST (12.2 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T182602Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048.json sha256 89ee35a7fb8e6ec8fd63aa7e9e07a5ed1c3448eed3d10cebf357887e810a94fe
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps212 tok/s {"prompt_tokens": [10787, 10703, 10755], "ttft_s": [50.583, 51.91, 50.672]}
ttft_prefill_32k_tps221 tok/s {"prompt_tokens": [42782, 42887, 42759], "ttft_s": [192.559, 193.833, 193.907]}
ttft_prefill_tps212 tok/s {"prompt_tokens": [10787, 10703, 10755], "ttft_s": [50.583, 51.91, 50.672]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitd4f093c724ebfd683feed823be35a11d4c009375 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T185415Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-ttft.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2

20261003T190629Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-stability

! fail stability T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2
Started2026-10-04 05:06 AEST
Finished2026-10-04 05:16 AEST (10.4 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 30 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 8.854 < floor 15.0

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T182602Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048.json sha256 89ee35a7fb8e6ec8fd63aa7e9e07a5ed1c3448eed3d10cebf357887e810a94fe
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max9,732 tokens
context.headroom_min121,298 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls30 calls {"per_min": 3.0, "requests": 21}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps8.85 tok/s
window 1: 9.19 tok/swindow 2: 9.0 tok/swindow 3: 8.95 tok/swindow 4: 8.92 tok/swindow 5: 9.14 tok/smin 8.92 · max 9.19 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 274, "min_at_generation_s": 232.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps9.09 tok/s {"generation_seconds": 393.673, "generated_tokens": 3578.8}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 578.9, "requests": 21, "tool_calls": 30, "completion_tokens": 3601, "occupied_max": 9732}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 21.1, "requests": 1, "tool_calls": 1, "completion_tokens": 0, "occupied_max": 0}] list
tasks_over_64k[] instance ids
completion_tokens3,601 tokens {"finish_reasons": {"tool_calls": 21}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 30 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 8.85 tok/s, whole run 9.09 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitd4f093c724ebfd683feed823be35a11d4c009375 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T190629Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-stability.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2

20261003T191656Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-lab

! fail lab T20b eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2
Started2026-10-04 05:16 AEST
Finished2026-10-04 05:31 AEST (15.0 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T182602Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048.json sha256 89ee35a7fb8e6ec8fd63aa7e9e07a5ed1c3448eed3d10cebf357887e810a94fe
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 16 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps9.50 tok/s
lab_prefill_tps199 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitd4f093c724ebfd683feed823be35a11d4c009375 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T191656Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2-lab.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2

20261003T193228Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-load

✓ pass load T22b eligibility: none

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2
Started2026-10-04 05:32 AEST
Finished2026-10-04 05:33 AEST (0.9 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! TTFT stability ✕

Launch

recorded by tools/run.py (artifacts/runs/20261003T193228Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4.json sha256 aa948a723db410fdb552afee2e3ac6e4d2fa58a7a7d510ddd629de6bd166b13a
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds51.9 s
host_ram_drop_gb115 GB
vram_ready_mib22,318 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit956afbaae18f13c44809e1d61a0a1568dca9db78 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T193228Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-load.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2

20261003T193321Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-lab

! fail lab T22b eligibility: none failed stage: gates

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2
Started2026-10-04 05:33 AEST
Finished2026-10-04 05:48 AEST (15.0 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! TTFT stability ✕

Launch

recorded by tools/run.py (artifacts/runs/20261003T193228Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4.json sha256 aa948a723db410fdb552afee2e3ac6e4d2fa58a7a7d510ddd629de6bd166b13a
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}
lab_decode_c1_tps8.60 tok/s
lab_prefill_tps147 tok/s "context-gate prompt tokens / total request seconds"
mtp_acceptance_rate0.5403 fraction {"draft_tokens": 2889, "accepted_tokens": 1561, "source": "server log delta"}
mtp_accepted_tps8.60 tok/s "output tok/s including accepted draft tokens"

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit956afbaae18f13c44809e1d61a0a1568dca9db78 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T193321Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-lab.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2

20261003T194821Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-ttft

✓ pass TTFT T22b eligibility: none

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2
Started2026-10-04 05:48 AEST
Finished2026-10-04 06:02 AEST (13.9 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! TTFT stability ✕

Launch

recorded by tools/run.py (artifacts/runs/20261003T193228Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4.json sha256 aa948a723db410fdb552afee2e3ac6e4d2fa58a7a7d510ddd629de6bd166b13a
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps199 tok/s {"prompt_tokens": [10721, 10696, 10751], "ttft_s": [53.479, 53.836, 54.249]}
ttft_prefill_32k_tps191 tok/s {"prompt_tokens": [42813, 42832, 42764], "ttft_s": [223.74, 224.908, 224.251]}
ttft_prefill_tps199 tok/s {"prompt_tokens": [10721, 10696, 10751], "ttft_s": [53.479, 53.836, 54.249]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit956afbaae18f13c44809e1d61a0a1568dca9db78 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T194821Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-ttft.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2

20261003T200216Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-stability

✕ crash stability T22b eligibility: none failed stage: measure

Configik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2
Started2026-10-04 06:02 AEST
Finished2026-10-04 06:12 AEST (10.1 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); 9 tool-call parse failure(s); decode floor not verifiable (too little generation)

Other runs of this config: load lab ! TTFT stability ✕

Launch

recorded by tools/run.py (artifacts/runs/20261003T193228Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4.json sha256 aa948a723db410fdb552afee2e3ac6e4d2fa58a7a7d510ddd629de6bd166b13a
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --cpu-moe --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 16 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --ctx-checkpoints 4
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max1,641 tokens
context.headroom_min129,356 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls0 calls {"per_min": 0.0, "requests": 9}
parse_failures9 responses {"wire": {"unparsed_markup": 9}, "harness_format_errors": {"no_tool_call": 9}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors9 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tpsunavailable only 58.214 s of generation; need > 120 s for one window
decode_whole_run_tps12.5 tok/s {"generation_seconds": 58.214, "generated_tokens": 729.8}
tasks_attempted5 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved0 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomyes {"upstream_connection_failures": 54, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 41.0, "requests": 3, "tool_calls": 0, "completion_tokens": 274, "occupied_max": 1559}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 36.8, "requests": 3, "tool_calls": 0, "completion_tokens": 248, "occupied_max": 1641}, {"attempt": 3, "round": 1, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 34.6, "requests": 3, "tool_calls": 0, "completion_tokens": 217, "occupied_max": 1609}, {"attempt": 4, "round": 1, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "Error:BadGatewayError: litellm.BadGatewayError: BadGatewayError: OpenAIException - Cannot connect to host 127.0.0.1:18080 ssl:default [Connect call failed ('127.0.0.1', 18080)]", "submitted": false, "resolved": false, "seconds": 341.4}, {"attempt": 5, "round": 1, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 146.2}] list
tasks_over_64k[] instance ids
completion_tokens739 tokens {"finish_reasons": {"stop": 9}}
mtp_accepted_tpsunavailable not produced (crash/measure)
mtp_acceptance_rateunavailable not produced (crash/measure)

Agentic summary: 600 s, 0 tool calls, 9 parse failures, 0 repetition hits, tasks solved 0/5, decode min window – tok/s, whole run 12.5 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-31
cpus_offlinenone
machine_checks_this_boot0
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit956afbaae18f13c44809e1d61a0a1568dca9db78 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T200216Z-ik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2-stability.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012

20261003T220620Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012-load

✓ pass load T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012
Started2026-10-04 08:06 AEST
Finished2026-10-04 08:07 AEST (0.7 min)
Exit statusNone
Notes–

Other runs of this config: load stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T220620Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit.json sha256 f7129eb274fa7a838d7a8928e9e72cec61e157259a237393b1f99b1945e722ab
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds39.9 s
host_ram_drop_gb3.35 GB
vram_ready_mib68,976 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit2e517d99a59b71bbf630b44ce6204f3b85717adb clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261003T220620Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012-load.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012

20261003T221725Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012-stability

! fail stability T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012
Started2026-10-04 08:17 AEST
Finished2026-10-04 08:19 AEST (1.9 min)
Exit status0
Notesstability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode floor not verifiable (too little generation); agent driver exited 1

Other runs of this config: load stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T220620Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit.json sha256 f7129eb274fa7a838d7a8928e9e72cec61e157259a237393b1f99b1945e722ab
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no successful request
context.headroom_minunavailable configured context unknown (set CONFIGURED_CTX)
duration_sunavailable driver produced no timing
tool_calls0 calls {"per_min": null, "requests": 0}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": null, "saved": "artifacts: proxy/repetition/"}
decode_window_min_tpsunavailable only 0.0 s of generation; need > 120 s for one window
decode_whole_run_tpsunavailable no generation recorded
tasks_attempted0 attempts {"unfinished_at_cap": 0, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solvedskipped no attempts
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[] list
tasks_over_64k[] instance ids
completion_tokens0 tokens {"finish_reasons": {}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: – s, 0 tool calls, 0 parse failures, 0 repetition hits, tasks solved –/0, decode min window – tok/s, whole run – tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitd9944c1c5f811326cc98ee2701073db1981ab9b7 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261003T221725Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012-stability.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012

20261003T222524Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-load

✓ pass load T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012
Started2026-10-04 08:25 AEST
Finished2026-10-04 08:27 AEST (1.7 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! lab ✕ TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T222524Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048.json sha256 ada3d8f510a24f0d5811d7af83862f218e4475f6cd2935003b2162f2c40e76ce
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds101 s
host_ram_drop_gb40.4 GB
vram_ready_mib68,496 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitcf7ad504f20ebd9fc05fbafb689f4809d92ed210 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T222524Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012

20261003T222706Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-lab

! fail lab T21 eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012
Started2026-10-04 08:27 AEST
Finished2026-10-04 08:51 AEST (24.6 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ✕ TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T222524Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048.json sha256 ada3d8f510a24f0d5811d7af83862f218e4475f6cd2935003b2162f2c40e76ce
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps12.3 tok/s
lab_prefill_tps158 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitcf7ad504f20ebd9fc05fbafb689f4809d92ed210 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T222706Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012

20261003T225145Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-ttft

✓ pass TTFT T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012
Started2026-10-04 08:51 AEST
Finished2026-10-04 09:07 AEST (15.5 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! lab ✕ TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T222524Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048.json sha256 ada3d8f510a24f0d5811d7af83862f218e4475f6cd2935003b2162f2c40e76ce
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps166 tok/s {"prompt_tokens": [10720, 10806, 10675], "ttft_s": [64.46, 64.694, 64.976]}
ttft_prefill_32k_tps175 tok/s {"prompt_tokens": [42882, 42861, 42885], "ttft_s": [244.773, 246.002, 245.262]}
ttft_prefill_tps166 tok/s {"prompt_tokens": [10720, 10806, 10675], "ttft_s": [64.46, 64.694, 64.976]}

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitcf7ad504f20ebd9fc05fbafb689f4809d92ed210 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T225145Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-ttft.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012

20261003T230715Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-stability

! fail stability T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012
Started2026-10-04 09:07 AEST
Finished2026-10-04 09:17 AEST (10.5 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 26 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 11.388 < floor 15.0

Other runs of this config: load lab ! lab ✕ TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T222524Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048.json sha256 ada3d8f510a24f0d5811d7af83862f218e4475f6cd2935003b2162f2c40e76ce
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max8,118 tokens
context.headroom_min122,865 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls26 calls {"per_min": 2.6, "requests": 15}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps11.4 tok/s
window 1: 11.57 tok/swindow 2: 11.47 tok/swindow 3: 11.64 tok/smin 11.47 · max 11.64 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 143, "min_at_generation_s": 109.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps11.7 tok/s {"generation_seconds": 262.271, "generated_tokens": 3060.1}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 364.3, "requests": 12, "tool_calls": 20, "completion_tokens": 2740, "occupied_max": 8118}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 235.7, "requests": 4, "tool_calls": 6, "completion_tokens": 336, "occupied_max": 2784}] list
tasks_over_64k[] instance ids
completion_tokens3,076 tokens {"finish_reasons": {"tool_calls": 15}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 26 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 11.4 tok/s, whole run 11.7 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitcf7ad504f20ebd9fc05fbafb689f4809d92ed210 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T230715Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-stability.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012

20261003T231744Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-lab

✕ crash lab T21 eligibility: none failed stage: harness

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012
Started2026-10-04 09:17 AEST
Finished2026-10-04 09:29 AEST (11.3 min)
Exit status1
Notes–

Other runs of this config: load lab ! lab ✕ TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T222524Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048.json sha256 ada3d8f510a24f0d5811d7af83862f218e4475f6cd2935003b2162f2c40e76ce
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
gates_passedunavailable not produced (crash/harness)
lab_decode_c1_tpsunavailable not produced (crash/harness)
lab_prefill_tpsunavailable not produced (crash/harness)
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitcf7ad504f20ebd9fc05fbafb689f4809d92ed210 clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T231744Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012

20261003T232915Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-load

✓ pass load T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012
Started2026-10-04 09:29 AEST
Finished2026-10-04 09:30 AEST (1.3 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T232915Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048.json sha256 4112d34e61cb2bb3861a59a7bf7cc1ce23cbba5e7f1b7cb3ce73f36dcfa2b7af
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds79.7 s
host_ram_drop_gb2.99 GB
vram_ready_mib67,148 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitd6382f80d3a94ab1a97311c8a27427f2b5c7f9cb clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T232915Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012

20261003T233035Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-lab

! fail lab T21 eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012
Started2026-10-04 09:30 AEST
Finished2026-10-04 10:12 AEST (42.0 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T232915Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048.json sha256 4112d34e61cb2bb3861a59a7bf7cc1ce23cbba5e7f1b7cb3ce73f36dcfa2b7af
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps10.5 tok/s
lab_prefill_tps126 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit94e598f979af99cb8e62f9dacf2a0f1b258ccf3e clean
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261003T233035Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012

20261004T001236Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-ttft

✓ pass TTFT T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012
Started2026-10-04 10:12 AEST
Finished2026-10-04 10:31 AEST (19.3 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T232915Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048.json sha256 4112d34e61cb2bb3861a59a7bf7cc1ce23cbba5e7f1b7cb3ce73f36dcfa2b7af
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps130 tok/s {"prompt_tokens": [10633, 10740, 10676], "ttft_s": [80.845, 98.056, 82.319]}
ttft_prefill_32k_tps143 tok/s {"prompt_tokens": [42717, 42962, 42901], "ttft_s": [298.369, 301.101, 295.768]}
ttft_prefill_tps130 tok/s {"prompt_tokens": [10633, 10740, 10676], "ttft_s": [80.845, 98.056, 82.319]}

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit376b9064b928086c6007ed16ab0d9267e111dd92 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T001236Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-ttft.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012

20261004T003153Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-stability

! fail stability T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012
Started2026-10-04 10:31 AEST
Finished2026-10-04 10:42 AEST (10.5 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 22 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 1 repetition hit(s); decode window min 9.855 < floor 15.0

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T232915Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048.json sha256 4112d34e61cb2bb3861a59a7bf7cc1ce23cbba5e7f1b7cb3ce73f36dcfa2b7af
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json

Metrics

context.configured131,072 tokens
context.occupied_max5,309 tokens
context.headroom_min125,652 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls22 calls {"per_min": 2.2, "requests": 16}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits1 hits {"ngram64x3": 1, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps9.86 tok/s
window 1: 10.15 tok/swindow 2: 10.0 tok/swindow 3: 10.01 tok/smin 10 · max 10.15 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 148, "min_at_generation_s": 195.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps10.1 tok/s {"generation_seconds": 267.587, "generated_tokens": 2697.7}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 225.2, "requests": 10, "tool_calls": 15, "completion_tokens": 1117, "occupied_max": 3970}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 374.8, "requests": 7, "tool_calls": 7, "completion_tokens": 1598, "occupied_max": 5309}] list
tasks_over_64k[] instance ids
completion_tokens2,715 tokens {"finish_reasons": {"tool_calls": 16}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 22 tool calls, 0 parse failures, 1 repetition hits, tasks solved 1/2, decode min window 9.86 tok/s, whole run 10.1 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit376b9064b928086c6007ed16ab0d9267e111dd92 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T003153Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-stability.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012

20261004T004223Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-lab

! fail lab T21 eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012
Started2026-10-04 10:42 AEST
Finished2026-10-04 11:09 AEST (27.2 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261003T232915Z-lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ3_XXS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048.json sha256 4112d34e61cb2bb3861a59a7bf7cc1ce23cbba5e7f1b7cb3ce73f36dcfa2b7af
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ3_XXS.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps10.4 tok/s
lab_prefill_tps132 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit376b9064b928086c6007ed16ab0d9267e111dd92 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T004223Z-llama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012

20261004T013855Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012-load

✓ pass load T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012
Started2026-10-04 11:38 AEST
Finished2026-10-04 11:39 AEST (0.9 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T013855Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048.json sha256 d46ba43c1169b3fd2e5d00ace10dc78cf7122918e1c259debec506f8a0766d22
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds55.0 s
host_ram_drop_gb3.83 GB
vram_ready_mib65,124 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit640a2b5606870f5adba5d1a5f8a557a075bbdecc dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T013855Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012

20261004T013950Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012-lab

! fail lab T21 eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012
Started2026-10-04 11:39 AEST
Finished2026-10-04 11:59 AEST (19.8 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T013855Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048.json sha256 d46ba43c1169b3fd2e5d00ace10dc78cf7122918e1c259debec506f8a0766d22
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps12.2 tok/s
lab_prefill_tps153 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit640a2b5606870f5adba5d1a5f8a557a075bbdecc dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T013950Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012

20261004T015940Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012-ttft

✓ pass TTFT T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012
Started2026-10-04 11:59 AEST
Finished2026-10-04 12:15 AEST (16.0 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T013855Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048.json sha256 d46ba43c1169b3fd2e5d00ace10dc78cf7122918e1c259debec506f8a0766d22
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps159 tok/s {"prompt_tokens": [10694, 10665, 10720], "ttft_s": [67.037, 66.91, 67.278]}
ttft_prefill_32k_tps169 tok/s {"prompt_tokens": [42746, 42816, 42870], "ttft_s": [254.159, 253.337, 253.851]}
ttft_prefill_tps159 tok/s {"prompt_tokens": [10694, 10665, 10720], "ttft_s": [67.037, 66.91, 67.278]}

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit640a2b5606870f5adba5d1a5f8a557a075bbdecc dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T015940Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012-ttft.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012

20261004T021546Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-load

✓ pass load T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012
Started2026-10-04 12:15 AEST
Finished2026-10-04 12:16 AEST (0.9 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T021546Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210.json sha256 1f7fcd2dbb10fca06661e2443bf4355bb609fd7677b4665a7891f97c2f975aa9
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds54.9 s
host_ram_drop_gb4.03 GB
vram_ready_mib65,124 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitfd89b522703b1628ea955d1581e800faae3e8e5f dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T021546Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012

20261004T021642Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-lab

! fail lab T21 eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012
Started2026-10-04 12:16 AEST
Finished2026-10-04 12:27 AEST (11.2 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T021546Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210.json sha256 1f7fcd2dbb10fca06661e2443bf4355bb609fd7677b4665a7891f97c2f975aa9
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps11.8 tok/s
lab_prefill_tps219 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitfd89b522703b1628ea955d1581e800faae3e8e5f dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T021642Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012

20261004T022756Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-ttft

✓ pass TTFT T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012
Started2026-10-04 12:27 AEST
Finished2026-10-04 12:39 AEST (11.1 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T021546Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210.json sha256 1f7fcd2dbb10fca06661e2443bf4355bb609fd7677b4665a7891f97c2f975aa9
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps240 tok/s {"prompt_tokens": [10656, 10730, 10756], "ttft_s": [44.441, 45.48, 44.34]}
ttft_prefill_32k_tps242 tok/s {"prompt_tokens": [42911, 42857, 42784], "ttft_s": [177.67, 176.719, 177.974]}
ttft_prefill_tps240 tok/s {"prompt_tokens": [10656, 10730, 10756], "ttft_s": [44.441, 45.48, 44.34]}

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitfd89b522703b1628ea955d1581e800faae3e8e5f dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T022756Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-ttft.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012

20261004T023904Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-stability

! fail stability T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012
Started2026-10-04 12:39 AEST
Finished2026-10-04 12:49 AEST (10.4 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 24 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 11.035 < floor 15.0

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T021546Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210.json sha256 1f7fcd2dbb10fca06661e2443bf4355bb609fd7677b4665a7891f97c2f975aa9
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max6,174 tokens
context.headroom_min124,753 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls24 calls {"per_min": 2.4, "requests": 16}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps11.0 tok/s
window 1: 11.25 tok/swindow 2: 11.07 tok/smin 11.07 · max 11.25 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 78, "min_at_generation_s": 109.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps11.3 tok/s {"generation_seconds": 197.184, "generated_tokens": 2222.9}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 266.1, "requests": 13, "tool_calls": 19, "completion_tokens": 2034, "occupied_max": 6174}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 333.9, "requests": 4, "tool_calls": 5, "completion_tokens": 206, "occupied_max": 3300}] list
tasks_over_64k[] instance ids
completion_tokens2,240 tokens {"finish_reasons": {"tool_calls": 16}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 24 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 11.0 tok/s, whole run 11.3 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitfd89b522703b1628ea955d1581e800faae3e8e5f dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T023904Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-stability.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012

20261004T024930Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-lab

! fail lab T21 eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012
Started2026-10-04 12:49 AEST
Finished2026-10-04 13:05 AEST (16.4 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T021546Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210.json sha256 1f7fcd2dbb10fca06661e2443bf4355bb609fd7677b4665a7891f97c2f975aa9
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps11.8 tok/s
lab_prefill_tps218 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitfd89b522703b1628ea955d1581e800faae3e8e5f dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T024930Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2

20261004T030555Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2-load

✓ pass load T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2
Started2026-10-04 13:05 AEST
Finished2026-10-04 13:06 AEST (0.8 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T030555Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15.json sha256 72b9b0b1ecb00a8819e5a83e80deef181225276cbaa232ea9d7ed944c7391289
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds45.9 s
host_ram_drop_gb3.52 GB
vram_ready_mib21,944 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit0ace6784b9e48161d6b2ab80af2c1887ace48181 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T030555Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2

20261004T030642Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2-lab

! fail lab T20b eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2
Started2026-10-04 13:06 AEST
Finished2026-10-04 13:29 AEST (22.3 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T030555Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15.json sha256 72b9b0b1ecb00a8819e5a83e80deef181225276cbaa232ea9d7ed944c7391289
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps9.50 tok/s
lab_prefill_tps201 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbc3d891210f23ffb6504b865b4216f82e08a42ef dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T030642Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2

20261004T032900Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2-ttft

✓ pass TTFT T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2
Started2026-10-04 13:29 AEST
Finished2026-10-04 13:41 AEST (12.2 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T030555Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15.json sha256 72b9b0b1ecb00a8819e5a83e80deef181225276cbaa232ea9d7ed944c7391289
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps213 tok/s {"prompt_tokens": [10652, 10738, 10639], "ttft_s": [50.031, 50.271, 51.458]}
ttft_prefill_32k_tps222 tok/s {"prompt_tokens": [42844, 42844, 42698], "ttft_s": [192.579, 195.095, 192.023]}
ttft_prefill_tps213 tok/s {"prompt_tokens": [10652, 10738, 10639], "ttft_s": [50.031, 50.271, 51.458]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbc3d891210f23ffb6504b865b4216f82e08a42ef dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T032900Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2-ttft.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2

20261004T034114Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-load

✓ pass load T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2
Started2026-10-04 13:41 AEST
Finished2026-10-04 13:42 AEST (0.8 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T034114Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096.json sha256 ee2f2a2da672b25768c3764ea94939d349c3e8c3616b4ebf5a18302b922972be
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds45.9 s
host_ram_drop_gb4.55 GB
vram_ready_mib22,914 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit2dbf31a09288732fa781855672b4d25178eb4682 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T034114Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2

20261004T034201Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-lab

! fail lab T20b eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2
Started2026-10-04 13:42 AEST
Finished2026-10-04 13:55 AEST (13.9 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T034114Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096.json sha256 ee2f2a2da672b25768c3764ea94939d349c3e8c3616b4ebf5a18302b922972be
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps9.60 tok/s
lab_prefill_tps288 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit2dbf31a09288732fa781855672b4d25178eb4682 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T034201Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2

20261004T035555Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-ttft

✓ pass TTFT T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2
Started2026-10-04 13:55 AEST
Finished2026-10-04 14:04 AEST (8.5 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T034114Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096.json sha256 ee2f2a2da672b25768c3764ea94939d349c3e8c3616b4ebf5a18302b922972be
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps321 tok/s {"prompt_tokens": [10699, 10766, 10755], "ttft_s": [33.318, 33.541, 34.552]}
ttft_prefill_32k_tps316 tok/s {"prompt_tokens": [42816, 42866, 42738], "ttft_s": [135.262, 135.596, 135.444]}
ttft_prefill_tps321 tok/s {"prompt_tokens": [10699, 10766, 10755], "ttft_s": [33.318, 33.541, 34.552]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit2dbf31a09288732fa781855672b4d25178eb4682 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T035555Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-ttft.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2

20261004T040423Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-stability

! fail stability T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2
Started2026-10-04 14:04 AEST
Finished2026-10-04 14:14 AEST (10.5 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 27 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 1 repetition hit(s); decode window min 8.847 < floor 15.0

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T034114Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096.json sha256 ee2f2a2da672b25768c3764ea94939d349c3e8c3616b4ebf5a18302b922972be
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max6,632 tokens
context.headroom_min124,337 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls27 calls {"per_min": 2.7, "requests": 17}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits1 hits {"ngram64x3": 1, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps8.85 tok/s
window 1: 9.1 tok/swindow 2: 8.89 tok/swindow 3: 9.14 tok/swindow 4: 9.3 tok/swindow 5: 9.15 tok/swindow 6: 9.01 tok/smin 8.89 · max 9.3 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 314, "min_at_generation_s": 123.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps9.10 tok/s {"generation_seconds": 433.213, "generated_tokens": 3940.0}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 331.4, "requests": 11, "tool_calls": 18, "completion_tokens": 2064, "occupied_max": 6088}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 268.6, "requests": 7, "tool_calls": 9, "completion_tokens": 1894, "occupied_max": 6632}] list
tasks_over_64k[] instance ids
completion_tokens3,958 tokens {"finish_reasons": {"tool_calls": 17}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 27 tool calls, 0 parse failures, 1 repetition hits, tasks solved 1/2, decode min window 8.85 tok/s, whole run 9.10 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit2dbf31a09288732fa781855672b4d25178eb4682 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T040423Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-stability.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2

20261004T041451Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-lab

! fail lab T20b eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2
Started2026-10-04 14:14 AEST
Finished2026-10-04 14:23 AEST (8.6 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T034114Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096.json sha256 ee2f2a2da672b25768c3764ea94939d349c3e8c3616b4ebf5a18302b922972be
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 4096 --ubatch-size 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps9.60 tok/s
lab_prefill_tps288 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit2dbf31a09288732fa781855672b4d25178eb4682 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T041451Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2-lab.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012

20261004T042329Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012-load

✕ crash load T22c eligibility: none failed stage: load

Configik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012
Started2026-10-04 14:23 AEST
Finished2026-10-04 14:24 AEST (0.7 min)
Exit status1
Notesche_init: CUDA2 KV buffer size = 570.81 MiB llama_init_from_model: KV self size = 748.00 MiB, c^KV (q8_0): 748.00 MiB, kv^T: not used llama_init_from_model: KV self indexer size = 704.00 MiB (f16) llama_init_from_model: CUDA_Host output buffer size = 0.59 MiB ggml_backend_cuda_buffer_type_alloc_buffer: allocating 3488.51 MiB on device 0: cudaMalloc failed: out of memory ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 3657965696 llama_init_from_model: failed to allocate compute buffers ~ggml_backend_cuda_context: have 0 graphs ~ggml_backend_cuda_context: have 0 graphs ~ggml_backend_cuda_context: have 0 graphs llama_init_from_gpt_params: error: failed to create context with model '/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf'

Launch

recorded by tools/run.py (artifacts/runs/20261004T042329Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210.json sha256 b1d4a0440963d486728825e444f7ac6736f7d39444e483cdf843d738cf41f923
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_secondsunavailable not produced (crash/load)
host_ram_drop_gbunavailable not produced (crash/load)
vram_ready_mibunavailable not produced (crash/load)

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit562ad23a6a0d69a48c21a7611e7c2d9f0984a1ed dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T042329Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012-load.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012

20261004T042414Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012-load

✕ crash load T22c eligibility: none failed stage: load

Configik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012
Started2026-10-04 14:24 AEST
Finished2026-10-04 14:25 AEST (0.8 min)
Exit status1
Notesche_init: CUDA2 KV buffer size = 570.81 MiB llama_init_from_model: KV self size = 816.00 MiB, c^KV (q8_0): 816.00 MiB, kv^T: not used llama_init_from_model: KV self indexer size = 768.00 MiB (f16) llama_init_from_model: CUDA_Host output buffer size = 0.61 MiB ggml_backend_cuda_buffer_type_alloc_buffer: allocating 3488.51 MiB on device 0: cudaMalloc failed: out of memory ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 3657965696 llama_init_from_model: failed to allocate compute buffers ~ggml_backend_cuda_context: have 0 graphs ~ggml_backend_cuda_context: have 0 graphs ~ggml_backend_cuda_context: have 0 graphs llama_init_from_gpt_params: error: failed to create context with model '/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf'

Launch

recorded by tools/run.py (artifacts/runs/20261004T042414Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210.json sha256 9f2b351bbcb5a9f58f9c86180a6023d2ce919c2149f9640d3778c925a0ec8881
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_secondsunavailable not produced (crash/load)
host_ram_drop_gbunavailable not produced (crash/load)
vram_ready_mibunavailable not produced (crash/load)

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit0c585c1d628c0c4a04e8994c51e5954b718fabea dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T042414Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012

20261004T042504Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-load

✓ pass load T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012
Started2026-10-04 14:25 AEST
Finished2026-10-04 14:26 AEST (1.5 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T042504Z-lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210- --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210.json sha256 6a17b74e230df72d3cf6d97693f004f278ead699c76a664b7f92b5fe36813807
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds91.2 s
host_ram_drop_gb3.05 GB
vram_ready_mib63,084 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit743ba9ddf00e0b7eb28b9180241f4a3fb9322a06 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T042504Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012

20261004T042635Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-lab

! fail lab T21 eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012
Started2026-10-04 14:26 AEST
Finished2026-10-04 14:52 AEST (25.7 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T042504Z-lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210- --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210.json sha256 6a17b74e230df72d3cf6d97693f004f278ead699c76a664b7f92b5fe36813807
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps8.00 tok/s
lab_prefill_tps139 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit743ba9ddf00e0b7eb28b9180241f4a3fb9322a06 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T042635Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012

20261004T045220Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-ttft

✓ pass TTFT T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012
Started2026-10-04 14:52 AEST
Finished2026-10-04 15:08 AEST (16.7 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T042504Z-lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210- --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210.json sha256 6a17b74e230df72d3cf6d97693f004f278ead699c76a664b7f92b5fe36813807
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps149 tok/s {"prompt_tokens": [10777, 10733, 10753], "ttft_s": [130.896, 71.944, 64.716]}
ttft_prefill_32k_tps176 tok/s {"prompt_tokens": [42666, 42838, 42755], "ttft_s": [243.017, 245.943, 242.684]}
ttft_prefill_tps149 tok/s {"prompt_tokens": [10777, 10733, 10753], "ttft_s": [130.896, 71.944, 64.716]}

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit743ba9ddf00e0b7eb28b9180241f4a3fb9322a06 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T045220Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-ttft.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012

20261004T050859Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-stability

! fail stability T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012
Started2026-10-04 15:08 AEST
Finished2026-10-04 15:19 AEST (10.5 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 25 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 7.807 < floor 15.0

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T042504Z-lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210- --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210.json sha256 6a17b74e230df72d3cf6d97693f004f278ead699c76a664b7f92b5fe36813807
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_max6,852 tokens
context.headroom_min123,939 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls25 calls {"per_min": 2.5, "requests": 19}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps7.81 tok/s
window 1: 7.91 tok/swindow 2: 7.96 tok/swindow 3: 8.06 tok/smin 7.91 · max 8.06 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 175, "min_at_generation_s": 107.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps7.95 tok/s {"generation_seconds": 294.515, "generated_tokens": 2340.5}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 277.5, "requests": 9, "tool_calls": 12, "completion_tokens": 1323, "occupied_max": 4597}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 322.5, "requests": 11, "tool_calls": 13, "completion_tokens": 1038, "occupied_max": 6852}] list
tasks_over_64k[] instance ids
completion_tokens2,361 tokens {"finish_reasons": {"tool_calls": 19}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 25 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 7.81 tok/s, whole run 7.95 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit743ba9ddf00e0b7eb28b9180241f4a3fb9322a06 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T050859Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-stability.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012

20261004T051928Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-lab

! fail lab T21 eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012
Started2026-10-04 15:19 AEST
Finished2026-10-04 15:38 AEST (19.0 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T042504Z-lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210- --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210.json sha256 6a17b74e230df72d3cf6d97693f004f278ead699c76a664b7f92b5fe36813807
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 2048 --ubatch-size 2048 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps8.30 tok/s
lab_prefill_tps156 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit743ba9ddf00e0b7eb28b9180241f4a3fb9322a06 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T051928Z-llama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012-lab.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012

20261004T053844Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-load

✓ pass load T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012
Started2026-10-04 15:38 AEST
Finished2026-10-04 15:39 AEST (0.7 min)
Exit statusNone
Notes–

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T053844Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210.json sha256 75dbaf82a0bd6bfa005b5ab1a0c00ed1eddb44fb5fe1060d7cbbd834d962f115
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds39.9 s
host_ram_drop_gb2.65 GB
vram_ready_mib65,510 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitfd8fc15971beec0df168ae2ee8c6c01aae2a7453 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T053844Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-load.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012

20261004T053925Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-lab

✓ pass lab T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012
Started2026-10-04 15:39 AEST
Finished2026-10-04 15:45 AEST (6.3 min)
Exit status0
Notesserved=/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T053844Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210.json sha256 75dbaf82a0bd6bfa005b5ab1a0c00ed1eddb44fb5fe1060d7cbbd834d962f115
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_max106,294 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps56.9 tok/s
lab_prefill_tps605 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commite17f05543581ac460774a08531a895566f2d209d dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T053925Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-lab.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012

20261004T054545Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-ttft

✓ pass TTFT T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012
Started2026-10-04 15:45 AEST
Finished2026-10-04 15:49 AEST (3.9 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T053844Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210.json sha256 75dbaf82a0bd6bfa005b5ab1a0c00ed1eddb44fb5fe1060d7cbbd834d962f115
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps987 tok/s {"prompt_tokens": [13709, 13794, 13718], "ttft_s": [14.16, 13.87, 13.896]}
ttft_prefill_32k_tps852 tok/s {"prompt_tokens": [54701, 54603, 54830], "ttft_s": [64.015, 64.121, 64.836]}
ttft_prefill_tps987 tok/s {"prompt_tokens": [13709, 13794, 13718], "ttft_s": [14.16, 13.87, 13.896]}

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commite17f05543581ac460774a08531a895566f2d209d dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T054545Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-ttft.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012

20261004T054940Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-stability

! fail stability T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012
Started2026-10-04 15:49 AEST
Finished2026-10-04 16:00 AEST (10.4 min)
Exit status0
Notesstability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 70 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 38.626 < floor 50.0

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T053844Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210.json sha256 75dbaf82a0bd6bfa005b5ab1a0c00ed1eddb44fb5fe1060d7cbbd834d962f115
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_max39,226 tokens
context.headroom_min91,731 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls70 calls {"per_min": 7.0, "requests": 54}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps38.6 tok/s
window 1: 48.84 tok/swindow 2: 45.07 tok/swindow 3: 41.96 tok/swindow 4: 39.7 tok/smin 39.7 · max 48.84 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 209, "min_at_generation_s": 267.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps45.4 tok/s {"generation_seconds": 328.55, "generated_tokens": 14907.4}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 49.5, "requests": 11, "tool_calls": 16, "completion_tokens": 2006, "occupied_max": 5874}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 550.5, "requests": 44, "tool_calls": 54, "completion_tokens": 12968, "occupied_max": 39226}] list
tasks_over_64k[] instance ids
completion_tokens14,974 tokens {"finish_reasons": {"tool_calls": 54}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 70 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 38.6 tok/s, whole run 45.4 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commite17f05543581ac460774a08531a895566f2d209d dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T054940Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-stability.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012

20261004T060007Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-lab

✓ pass lab T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012
Started2026-10-04 16:00 AEST
Finished2026-10-04 16:13 AEST (13.3 min)
Exit status0
Notesserved=/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T053844Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210.json sha256 75dbaf82a0bd6bfa005b5ab1a0c00ed1eddb44fb5fe1060d7cbbd834d962f115
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_max106,294 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps57.1 tok/s
lab_prefill_tps608 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commite17f05543581ac460774a08531a895566f2d209d dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T060007Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2

20261004T061329Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2-load

✓ pass load T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2
Started2026-10-04 16:13 AEST
Finished2026-10-04 16:14 AEST (0.8 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T061329Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 8192 --ubatch-size 8192
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192.json sha256 6a9c08474b62518d1e1c1d23f6b6b1f241e7a0e05c5faa5f176e3700eec9fefe
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 8192 --ubatch-size 8192
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds48.9 s
host_ram_drop_gb5.84 GB
vram_ready_mib22,186 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit5af5149471babca3e64d893eeea35a3a1a43fe06 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T061329Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2

20261004T061418Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2-lab

! fail lab T20b eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2
Started2026-10-04 16:14 AEST
Finished2026-10-04 16:23 AEST (9.0 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T061329Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 8192 --ubatch-size 8192
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192.json sha256 6a9c08474b62518d1e1c1d23f6b6b1f241e7a0e05c5faa5f176e3700eec9fefe
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 8192 --ubatch-size 8192
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps9.10 tok/s
lab_prefill_tps370 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitdd520bfd96d949a57962b56ed6cd8dd55f834916 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T061418Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2

20261004T062318Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2-ttft

✓ pass TTFT T20b eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2
Started2026-10-04 16:23 AEST
Finished2026-10-04 16:30 AEST (6.7 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T061329Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 8192 --ubatch-size 8192
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192.json sha256 6a9c08474b62518d1e1c1d23f6b6b1f241e7a0e05c5faa5f176e3700eec9fefe
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --batch-size 8192 --ubatch-size 8192
envCUDA_DEVICE_ORDER=PCI_BUS_ID
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps384 tok/s {"prompt_tokens": [10746, 10734, 10753], "ttft_s": [27.965, 27.98, 27.965]}
ttft_prefill_32k_tps404 tok/s {"prompt_tokens": [42688, 42881, 42819], "ttft_s": [105.798, 105.946, 106.334]}
ttft_prefill_tps384 tok/s {"prompt_tokens": [10746, 10734, 10753], "ttft_s": [27.965, 27.98, 27.965]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitdd520bfd96d949a57962b56ed6cd8dd55f834916 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T062318Z-llama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2-ttft.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012

20261004T063003Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012-load

✓ pass load T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012
Started2026-10-04 16:30 AEST
Finished2026-10-04 16:30 AEST (0.9 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T063003Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-g/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 4096 --ubatch-size 4096 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210.json sha256 e0b729222e80461545b33c6e10b8ff1d8ecaaf2fde1596824ce2a3e98b0ce76c
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 4096 --ubatch-size 4096 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds54.9 s
host_ram_drop_gb4.93 GB
vram_ready_mib65,688 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9a178e2fe77cc647e9542e1742a52c38931f07e1 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T063003Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012-load.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012

20261004T063058Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012-lab

! fail lab T21 eligibility: none failed stage: gates

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012
Started2026-10-04 16:30 AEST
Finished2026-10-04 16:48 AEST (17.9 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T063003Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-g/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 4096 --ubatch-size 4096 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210.json sha256 e0b729222e80461545b33c6e10b8ff1d8ecaaf2fde1596824ce2a3e98b0ce76c
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 4096 --ubatch-size 4096 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps11.2 tok/s
lab_prefill_tps258 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9a178e2fe77cc647e9542e1742a52c38931f07e1 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T063058Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012-lab.json

Runs / GLM-5.3-Flash / llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012

20261004T064850Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012-ttft

✓ pass TTFT T21 eligibility: none

Configllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012
Started2026-10-04 16:48 AEST
Finished2026-10-04 16:58 AEST (9.3 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T063003Z-lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-g/docker-run.txt)
docker run -d --name lab-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-g --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 4096 --ubatch-size 4096 --fit-target 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210.json sha256 e0b729222e80461545b33c6e10b8ff1d8ecaaf2fde1596824ce2a3e98b0ce76c
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --batch-size 4096 --ubatch-size 4096 --fit-target 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps298 tok/s {"prompt_tokens": [10829, 10677, 10700], "ttft_s": [37.303, 35.77, 35.957]}
ttft_prefill_32k_tps288 tok/s {"prompt_tokens": [42788, 42675, 42983], "ttft_s": [148.453, 148.434, 149.751]}
ttft_prefill_tps298 tok/s {"prompt_tokens": [10829, 10677, 10700], "ttft_s": [37.303, 35.77, 35.957]}

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9a178e2fe77cc647e9542e1742a52c38931f07e1 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T064850Z-llama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012-ttft.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012

20261004T065809Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-load

✓ pass load T22c eligibility: none

Configik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012
Started2026-10-04 16:58 AEST
Finished2026-10-04 16:58 AEST (0.7 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T065809Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096.json sha256 fe1c037653b578219259df5ac2f7fce9deed7442d4759cd749f0001ffe3df341
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds44.3 s
host_ram_drop_gb65.5 GB
vram_ready_mib68,156 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit981b077ba6f48c9b94c9bc0a253f95413b94b11e dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T065809Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-load.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012

20261004T065853Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-lab

! fail lab T22c eligibility: none failed stage: gates

Configik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012
Started2026-10-04 16:58 AEST
Finished2026-10-04 17:10 AEST (12.0 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T065809Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096.json sha256 fe1c037653b578219259df5ac2f7fce9deed7442d4759cd749f0001ffe3df341
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}
lab_decode_c1_tps14.0 tok/s
lab_prefill_tps185 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit981b077ba6f48c9b94c9bc0a253f95413b94b11e dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T065853Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-lab.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012

20261004T071053Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-ttft

✓ pass TTFT T22c eligibility: none

Configik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012
Started2026-10-04 17:10 AEST
Finished2026-10-04 17:21 AEST (10.4 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T065809Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096.json sha256 fe1c037653b578219259df5ac2f7fce9deed7442d4759cd749f0001ffe3df341
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps283 tok/s {"prompt_tokens": [10728, 10711, 10734], "ttft_s": [37.814, 37.788, 37.935]}
ttft_prefill_32k_tps253 tok/s {"prompt_tokens": [42827, 42804, 42935], "ttft_s": [167.744, 169.791, 169.931]}
ttft_prefill_tps283 tok/s {"prompt_tokens": [10728, 10711, 10734], "ttft_s": [37.814, 37.788, 37.935]}

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit981b077ba6f48c9b94c9bc0a253f95413b94b11e dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T071053Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-ttft.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012

20261004T072115Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-stability

! fail stability T22c eligibility: none

Configik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012
Started2026-10-04 17:21 AEST
Finished2026-10-04 17:31 AEST (10.1 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 68 tool-call parse failure(s); decode window min 13.255 < floor 15.0

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T065809Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096.json sha256 fe1c037653b578219259df5ac2f7fce9deed7442d4759cd749f0001ffe3df341
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max1,874 tokens
context.headroom_min129,132 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls0 calls {"per_min": 0.0, "requests": 68}
parse_failures68 responses {"wire": {"unparsed_markup": 68}, "harness_format_errors": {"no_tool_call": 68}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors68 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps13.3 tok/s
window 1: 13.31 tok/swindow 2: 13.29 tok/swindow 3: 13.29 tok/swindow 4: 13.35 tok/swindow 5: 13.34 tok/smin 13.29 · max 13.35 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 299, "min_at_generation_s": 112.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps13.3 tok/s {"generation_seconds": 418.663, "generated_tokens": 5576.2}
tasks_attempted23 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved0 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 30.9, "requests": 3, "tool_calls": 0, "completion_tokens": 246, "occupied_max": 1559}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 27.9, "requests": 3, "tool_calls": 0, "completion_tokens": 217, "occupied_max": 1641}, {"attempt": 3, "round": 1, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 37.7, "requests": 3, "tool_calls": 0, "completion_tokens": 348, "occupied_max": 1609}, {"attempt": 4, "round": 1, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 31.6, "requests": 3, "tool_calls": 0, "completion_tokens": 263, "occupied_max": 1874}, {"attempt": 5, "round": 1, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 31.2, "requests": 3, "tool_calls": 0, "completion_tokens": 260, "occupied_max": 1770}, {"attempt": 6, "round": 1, "instance_id": "astropy__astropy-12907", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 25.1, "requests": 3, "tool_calls": 0, "completion_tokens": 181, "occupied_max": 1716}, {"attempt": 7, "round": 1, "instance_id": "pytest-dev__pytest-7373", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 27.6, "requests": 3, "tool_calls": 0, "completion_tokens": 214, "occupied_max": 1641}, {"attempt": 8, "round": 1, "instance_id": "sympy__sympy-20590", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 39.7, "requests": 3, "tool_calls": 0, "completion_tokens": 375, "occupied_max": 1578}, {"attempt": 9, "round": 2, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 23.2, "requests": 3, "tool_calls": 0, "completion_tokens": 240, "occupied_max": 1559}, {"attempt": 10, "round": 2, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 21.1, "requests": 3, "tool_calls": 0, "completion_tokens": 212, "occupied_max": 1641}, {"attempt": 11, "round": 2, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 26.7, "requests": 3, "tool_calls": 0, "completion_tokens": 286, "occupied_max": 1609}, {"attempt": 12, "round": 2, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 25.0, "requests": 3, "tool_calls": 0, "completion_tokens": 262, "occupied_max": 1874}, {"attempt": 13, "round": 2, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 24.6, "requests": 3, "tool_calls": 0, "completion_tokens": 257, "occupied_max": 1770}, {"attempt": 14, "round": 2, "instance_id": "astropy__astropy-12907", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 19.0, "requests": 3, "tool_calls": 0, "completion_tokens": 184, "occupied_max": 1716}, {"attempt": 15, "round": 2, "instance_id": "pytest-dev__pytest-7373", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 20.8, "requests": 3, "tool_calls": 0, "completion_tokens": 208, "occupied_max": 1641}, {"attempt": 16, "round": 2, "instance_id": "sympy__sympy-20590", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 29.2, "requests": 3, "tool_calls": 0, "completion_tokens": 319, "occupied_max": 1578}, {"attempt": 17, "round": 3, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 23.2, "requests": 3, "tool_calls": 0, "completion_tokens": 240, "occupied_max": 1559}, {"attempt": 18, "round": 3, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 21.1, "requests": 3, "tool_calls": 0, "completion_tokens": 212, "occupied_max": 1641}, {"attempt": 19, "round": 3, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 26.8, "requests": 3, "tool_calls": 0, "completion_tokens": 286, "occupied_max": 1609}, {"attempt": 20, "round": 3, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 25.0, "requests": 3, "tool_calls": 0, "completion_tokens": 262, "occupied_max": 1874}, {"attempt": 21, "round": 3, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 24.6, "requests": 3, "tool_calls": 0, "completion_tokens": 257, "occupied_max": 1770}, {"attempt": 22, "round": 3, "instance_id": "astropy__astropy-12907", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 19.0, "requests": 3, "tool_calls": 0, "completion_tokens": 184, "occupied_max": 1716}, {"attempt": 23, "round": 3, "instance_id": "pytest-dev__pytest-7373", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 19.0, "requests": 3, "tool_calls": 0, "completion_tokens": 133, "occupied_max": 1522}] list
tasks_over_64k[] instance ids
completion_tokens5,646 tokens {"finish_reasons": {"stop": 68}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 0 tool calls, 68 parse failures, 0 repetition hits, tasks solved 0/23, decode min window 13.3 tok/s, whole run 13.3 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit981b077ba6f48c9b94c9bc0a253f95413b94b11e dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T072115Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-stability.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012

20261004T073119Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-lab

! fail lab T22c eligibility: none failed stage: gates

Configik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012
Started2026-10-04 17:31 AEST
Finished2026-10-04 17:42 AEST (11.2 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! lab ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T065809Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096.json sha256 fe1c037653b578219259df5ac2f7fce9deed7442d4759cd749f0001ffe3df341
MTP / speculativeoff
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --fit-margin 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}
lab_decode_c1_tps13.4 tok/s
lab_prefill_tps189 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit981b077ba6f48c9b94c9bc0a253f95413b94b11e dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T073119Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012-lab.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012

20261004T074233Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012-load

✕ crash load T22c eligibility: none failed stage: load

Configik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012
Started2026-10-04 17:42 AEST
Finished2026-10-04 17:43 AEST (0.9 min)
Exit status1
Notes c^KV (q8_0): 68.00 MiB, kv^T: not used llama_init_from_model: KV self indexer size = 64.00 MiB (f16) llama_init_from_model: CUDA_Host output buffer size = 0.61 MiB ggml_backend_cuda_buffer_type_alloc_buffer: allocating 2754.25 MiB on device 0: cudaMalloc failed: out of memory ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 2888039424 llama_init_from_model: failed to allocate compute buffers ~ggml_backend_cuda_context: have 0 graphs ~ggml_backend_cuda_context: have 0 graphs ~ggml_backend_cuda_context: have 0 graphs common_speculative_init: failed to create MTP context srv init: failed to initialize recurrent speculative context ~ggml_backend_cuda_context: have 13 graphs ~ggml_backend_cuda_context: have 14 graphs ~ggml_backend_cuda_context: have 11 graphs

Launch

recorded by tools/run.py (artifacts/runs/20261004T074233Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 4096
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096.json sha256 d21078e38ec9e96da581e9003a2aaae041591909413dbf70738357df0829341e
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_secondsunavailable not produced (crash/load)
host_ram_drop_gbunavailable not produced (crash/load)
vram_ready_mibunavailable not produced (crash/load)

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitf3cbbe1b98ba109e9c4abaac8d34bc7b6e2775ae dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T074233Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012-load.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012

20261004T080048Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012-load

✓ pass load T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012
Started2026-10-04 18:00 AEST
Finished2026-10-04 18:01 AEST (0.7 min)
Exit statusNone
Notes–

Other runs of this config: load lab TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T080048Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048.json sha256 e882737bdeeadb764f93582d803f25626630b8b0a09fb20716d71d1c2c7be516
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds39.9 s
host_ram_drop_gb3.74 GB
vram_ready_mib65,396 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit65a204856ba6e28ca395fa6aab47a88e41c62a1a dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T080048Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012-load.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012

20261004T080128Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012-lab

✓ pass lab T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012
Started2026-10-04 18:01 AEST
Finished2026-10-04 18:06 AEST (5.5 min)
Exit status0
Notesserved=/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf

Other runs of this config: load lab TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T080048Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048.json sha256 e882737bdeeadb764f93582d803f25626630b8b0a09fb20716d71d1c2c7be516
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_max106,294 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps44.3 tok/s
lab_prefill_tps717 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit65a204856ba6e28ca395fa6aab47a88e41c62a1a dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T080128Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012-lab.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012

20261004T080657Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012-ttft

✓ pass TTFT T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012
Started2026-10-04 18:06 AEST
Finished2026-10-04 18:10 AEST (3.6 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T080048Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 2048 --ubatch-size 2048
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048.json sha256 e882737bdeeadb764f93582d803f25626630b8b0a09fb20716d71d1c2c7be516
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 2048 --ubatch-size 2048
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps1,067 tok/s {"prompt_tokens": [13695, 13734, 13847], "ttft_s": [12.963, 12.872, 12.965]}
ttft_prefill_32k_tps931 tok/s {"prompt_tokens": [54728, 54685, 54604], "ttft_s": [58.434, 58.71, 58.906]}
ttft_prefill_tps1,067 tok/s {"prompt_tokens": [13695, 13734, 13847], "ttft_s": [12.963, 12.872, 12.965]}

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit65a204856ba6e28ca395fa6aab47a88e41c62a1a dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T080657Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub2048-gpu012-ttft.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012

20261004T081034Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012-load

✓ pass load T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012
Started2026-10-04 18:10 AEST
Finished2026-10-04 18:11 AEST (0.7 min)
Exit statusNone
Notes–

Other runs of this config: load lab TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T081034Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 4096 --ubatch-size 4096
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096.json sha256 bb5c0e11a6559307e48f7f8f285d3439f9c7872614fa95ee033b2d48ca370dbd
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 4096 --ubatch-size 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds39.9 s
host_ram_drop_gb4.34 GB
vram_ready_mib65,508 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbfcf5657ea4ed56f452dc532e3311cf77750c3d9 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T081034Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012-load.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012

20261004T081114Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012-lab

✓ pass lab T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012
Started2026-10-04 18:11 AEST
Finished2026-10-04 18:19 AEST (8.1 min)
Exit status0
Notesserved=/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf

Other runs of this config: load lab TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T081034Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 4096 --ubatch-size 4096
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096.json sha256 bb5c0e11a6559307e48f7f8f285d3439f9c7872614fa95ee033b2d48ca370dbd
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 4096 --ubatch-size 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_max106,294 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps35.0 tok/s
lab_prefill_tps713 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbfcf5657ea4ed56f452dc532e3311cf77750c3d9 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T081114Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012-lab.json

Runs / Qwen3.8-Flash-Next / llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012

20261004T081919Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012-ttft

✓ pass TTFT T11 eligibility: none

Configllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012
Started2026-10-04 18:19 AEST
Finished2026-10-04 18:23 AEST (3.8 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab TTFT

Launch

recorded by tools/run.py (artifacts/runs/20261004T081034Z-lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210/docker-run.txt)
docker run -d --name lab-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-Qwen3.8-Flash-Next-IQ4_XS:/models:ro --entrypoint /app/llama-server ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5 --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 4096 --ubatch-size 4096
Imageghcr.io/ggml-org/llama.cpp:server-cuda12-b11312@sha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Digestsha256:6cdf9529493b9581c421fec5628dcd4b8dd7c09507fc7d27ad739cc9205d66f5
Enginellama.cpp
Engine source–
Launch filelaunches/llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096.json sha256 bb5c0e11a6559307e48f7f8f285d3439f9c7872614fa95ee033b2d48ca370dbd
MTP / speculativeoff
argv/app/llama-server --model /models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf --ctx-size 131072 --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --parallel 1 --threads 15 --jinja --no-mmproj --metrics --host 0.0.0.0 --port 8080 --split-mode layer --fit-target 2048 --batch-size 4096 --ubatch-size 4096
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f · manifest models/bartowski-Qwen3.8-Flash-Next-IQ4_XS.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps980 tok/s {"prompt_tokens": [13680, 13707, 13810], "ttft_s": [14.13, 13.965, 14.096]}
ttft_prefill_32k_tps881 tok/s {"prompt_tokens": [54557, 54627, 54737], "ttft_s": [61.779, 62.069, 62.105]}
ttft_prefill_tps980 tok/s {"prompt_tokens": [13680, 13707, 13810], "ttft_s": [14.13, 13.965, 14.096]}

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbfcf5657ea4ed56f452dc532e3311cf77750c3d9 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T081919Z-llama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-ub4096-gpu012-ttft.json

Runs / none

20261004T091142Z-hw-bw-gpu1-current-hw_bw

✓ pass hw bw T05a eligibility: none

Confighw-bw-gpu1-current
Started2026-10-04 19:11 AEST
Finished2026-10-04 19:11 AEST (0.2 min)
Exit status0
Notes1 GiB pinned transfers; full table in detail

Launch

no container: static measurement

Metrics

context.configured–
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
h2d_gbps13.5 GB/s {"1MB_pin": 13.14, "1MB_page": 13.01, "4MB_pin": 13.39, "4MB_page": 13.12, "16MB_pin": 13.45, "16MB_page": 13.18, "64MB_pin": 13.47, "64MB_page": 13.18, "256MB_pin": 13.47, "256MB_page": 13.07, "1024MB_pin": 13.47, "1024MB_page": 13.12}
d2h_gbps13.2 GB/s {"1MB_pin": 12.95, "1MB_page": 6.3, "4MB_pin": 13.14, "4MB_page": 10.05, "16MB_pin": 13.2, "16MB_page": 11.16, "64MB_pin": 13.21, "64MB_page": 11.79, "256MB_pin": 13.21, "256MB_page": 12.04, "1024MB_pin": 13.22, "1024MB_page": 12.05}
host_memcpy_gbps16.6 GB/s

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit00709a719d4c6057fbeb8f52eadd20c2bbe2f46d dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/none/20261004T091142Z-hw-bw-gpu1-current-hw_bw.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012

20261004T091154Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-load

✓ pass load T22c eligibility: none

Configik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012
Started2026-10-04 19:11 AEST
Finished2026-10-04 19:12 AEST (0.8 min)
Exit statusNone
Notes–

Other runs of this config: load lab ! TTFT stability ✕

Launch

recorded by tools/run.py (artifacts/runs/20261004T091154Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168.json sha256 cf6116ffb2e3698333ab48e1435f99a6f037b39308f84e62e4fd24888017f06c
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds48.6 s
host_ram_drop_gb83.3 GB
vram_ready_mib64,680 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitc36c7e58741670e5e869f6857478b1e198932e9a dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T091154Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-load.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012

20261004T091243Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-lab

! fail lab T22c eligibility: none failed stage: gates

Configik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012
Started2026-10-04 19:12 AEST
Finished2026-10-04 19:26 AEST (13.5 min)
Exit status0
Notesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf

Other runs of this config: load lab ! TTFT stability ✕

Launch

recorded by tools/run.py (artifacts/runs/20261004T091154Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168.json sha256 cf6116ffb2e3698333ab48e1435f99a6f037b39308f84e62e4fd24888017f06c
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": false, "context": true, "speed": false}
lab_decode_c1_tps10.2 tok/s
lab_prefill_tps166 tok/s "context-gate prompt tokens / total request seconds"
mtp_acceptance_rate0.5137 fraction {"draft_tokens": 3366, "accepted_tokens": 1729, "source": "server log delta"}
mtp_accepted_tps10.2 tok/s "output tok/s including accepted draft tokens"

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit151350a4d65da543200768e08741f6eb5409c3ed dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T091243Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-lab.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012

20261004T092615Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-ttft

✓ pass TTFT T22c eligibility: none

Configik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012
Started2026-10-04 19:26 AEST
Finished2026-10-04 19:38 AEST (11.7 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab ! TTFT stability ✕

Launch

recorded by tools/run.py (artifacts/runs/20261004T091154Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168.json sha256 cf6116ffb2e3698333ab48e1435f99a6f037b39308f84e62e4fd24888017f06c
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps248 tok/s {"prompt_tokens": [10798, 10674, 10665], "ttft_s": [43.785, 43.033, 42.927]}
ttft_prefill_32k_tps224 tok/s {"prompt_tokens": [42970, 42867, 42822], "ttft_s": [191.654, 191.974, 191.365]}
ttft_prefill_tps248 tok/s {"prompt_tokens": [10798, 10674, 10665], "ttft_s": [43.785, 43.033, 42.927]}

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit42d955c4fa51a99398f09218c4a4c133bcf18938 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T092615Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-ttft.json

Runs / GLM-5.3-Flash / ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012

20261004T093800Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-stability

✕ crash stability T22c eligibility: none failed stage: measure

Configik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012
Started2026-10-04 19:38 AEST
Finished2026-10-04 19:48 AEST (10.1 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); 9 tool-call parse failure(s); decode floor not verifiable (too little generation)

Other runs of this config: load lab ! TTFT stability ✕

Launch

recorded by tools/run.py (artifacts/runs/20261004T091154Z-lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012/docker-run.txt)
docker run -d --name lab-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012 --security-opt no-new-privileges --gpus '"device=GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d,GPU-c67ac872-3371-a88c-aa92-d969daf6405e,GPU-76d3c6af-7f18-69d2-3145-899103de1722"' -p 127.0.0.1:18080:8080 --shm-size 16g -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e CUDA_VISIBLE_DEVICES=2,1,0 -v ~/models/local-ai-exp/bartowski-GLM-5.3-Flash-IQ2_M:/models:ro --entrypoint /app/llama-server local-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1 --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168
Imagelocal-ai-bench/ik-llama-cuda@sha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Digestsha256:d30fa2796c72509c9158a6202fed46c8ce159f3344f3735120be65a73a8137b1
Engineik_llama.cpp
Engine sourcehttps://github.com/ikawrakow/ik_llama.cpp/tree/5bf8f0fe4db98e9560de0f495232f8d3a733c678
Launch filelaunches/ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168.json sha256 cf6116ffb2e3698333ab48e1435f99a6f037b39308f84e62e4fd24888017f06c
MTP / speculative{"type": "mtp", "draft": "embedded NextN/MTP layer of the target GGUF (Q4_0 in bartowski files)", "spec_type": "mtp:n_max=3,p_min=0.0", "n_max": 3, "p_min": 0.0, "heads": 1, "ref": "ik_llama.cpp PR #2548 (n_max=3 best in author's n_max 3/4 comparison)"}
argv/app/llama-server --model /models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf --ctx-size 131072 --flash-attn on --mla-use 3 --dsa --cache-type-k q8_0 --n-gpu-layers 999 --fit --split-mode layer --batch-size 2048 --ubatch-size 2048 --parallel 1 --threads 15 --jinja --metrics --host 0.0.0.0 --port 8080 --spec-type mtp:n_max=3,p_min=0.0 --fit-margin 7168
envCUDA_DEVICE_ORDER=PCI_BUS_ID
CUDA_VISIBLE_DEVICES=2,1,0
bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 · manifest models/bartowski-GLM-5.3-Flash-IQ2_M.json

Metrics

context.configured131,072 tokens
context.occupied_max1,641 tokens
context.headroom_min129,336 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls0 calls {"per_min": 0.0, "requests": 9}
parse_failures9 responses {"wire": {"unparsed_markup": 9}, "harness_format_errors": {"no_tool_call": 9}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors9 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tpsunavailable only 41.172 s of generation; need > 120 s for one window
decode_whole_run_tps15.6 tok/s {"generation_seconds": 41.172, "generated_tokens": 642.7}
tasks_attempted5 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved0 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomyes {"upstream_connection_failures": 54, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 26.8, "requests": 3, "tool_calls": 0, "completion_tokens": 208, "occupied_max": 1559}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 28.3, "requests": 3, "tool_calls": 0, "completion_tokens": 227, "occupied_max": 1641}, {"attempt": 3, "round": 1, "instance_id": "psf__requests-2317", "exit_status": "RepeatedFormatError", "submitted": false, "resolved": false, "seconds": 27.1, "requests": 3, "tool_calls": 0, "completion_tokens": 217, "occupied_max": 1609}, {"attempt": 4, "round": 1, "instance_id": "matplotlib__matplotlib-23299", "exit_status": "Error:BadGatewayError: litellm.BadGatewayError: BadGatewayError: OpenAIException - Cannot connect to host 127.0.0.1:18080 ssl:default [Connect call failed ('127.0.0.1', 18080)]", "submitted": false, "resolved": false, "seconds": 343.6}, {"attempt": 5, "round": 1, "instance_id": "scikit-learn__scikit-learn-13439", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 174.2}] list
tasks_over_64k[] instance ids
completion_tokens652 tokens {"finish_reasons": {"stop": 9}}
mtp_accepted_tpsunavailable not produced (crash/measure)
mtp_acceptance_rateunavailable not produced (crash/measure)

Agentic summary: 600 s, 0 tool calls, 9 parse failures, 0 repetition hits, tasks solved 0/5, decode min window – tok/s, whole run 15.6 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 3 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0●NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1●NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit42d955c4fa51a99398f09218c4a4c133bcf18938 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T093800Z-ik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012-stability.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2

20261004T101015Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72

✓ pass load T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
Started2026-10-04 20:10 AEST
Finished2026-10-04 20:14 AEST (3.9 min)
Exit statusNone
Notes–

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Digestsha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds236 s
host_ram_drop_gb63.5 GB
vram_ready_mib20,818 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbdd5d466ac35145adc76b9f18538d7ef1373acce dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T101015Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2

20261004T101411Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72

✓ pass lab T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
Started2026-10-04 20:14 AEST
Finished2026-10-04 20:19 AEST (5.4 min)
Exit status0
Notesserved=flashnext

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Digestsha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max166,667 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps55.6 tok/s
lab_prefill_tps2,433 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbdd5d466ac35145adc76b9f18538d7ef1373acce dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T101411Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2

20261004T101937Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72

✓ pass kit sweep T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
Started2026-10-04 20:19 AEST
Finished2026-10-04 20:34 AEST (15.2 min)
Exit status0
Notessweep status: DONE

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Digestsha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
prefill_8k_tps2,289 tok/s {"min": 2269.6, "max": 2294.4, "n": 3}
prefill_16k_tps2,240 tok/s {"min": 2233.0, "max": 2247.3, "n": 3}
prefill_32k_tps2,678 tok/s {"min": 2672.6, "max": 2681.2, "n": 3}
prefill_64k_tps2,653 tok/s {"min": 2645.1, "max": 2655.6, "n": 3}
decode_c1_tps48.2 tok/s {"min": 47.55, "max": 48.85, "power_w": null}
decode_c2_tps53.1 tok/s {"min": 51.57, "max": 54.71, "power_w": null}
decode_c3_tps55.9 tok/s {"min": 55.26, "max": 56.61, "power_w": null}
decode_c4_tps53.3 tok/s {"min": 51.51, "max": 55.06, "power_w": null}
decode_c1_32k_tps49.9 tok/s {"min": 49.58, "max": 50.2, "power_w": null}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9c490cfb769bb4bce8cd979c9e96a84128be7bc8 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T101937Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2

20261004T103452Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72

✓ pass TTFT T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
Started2026-10-04 20:34 AEST
Finished2026-10-04 20:36 AEST (1.4 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Digestsha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps2,017 tok/s {"prompt_tokens": [13801, 13745, 13788], "ttft_s": [6.842, 6.83, 6.822]}
ttft_prefill_32k_tps2,687 tok/s {"prompt_tokens": [54693, 54557, 54874], "ttft_s": [20.353, 20.312, 20.371]}
ttft_prefill_tps2,017 tok/s {"prompt_tokens": [13801, 13745, 13788], "ttft_s": [6.842, 6.83, 6.822]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9c490cfb769bb4bce8cd979c9e96a84128be7bc8 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T103452Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2

20261004T103614Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72

✓ pass quality panel T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
Started2026-10-04 20:36 AEST
Finished2026-10-04 20:36 AEST (0.4 min)
Exit status0
Notespanel=qwen3.8-flash-next-exl3-ref-panel.json rc=0

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Digestsha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
top1_agreement0.98867 fraction
mean_kl0.00096647 nats
in_bandyes "top1>=0.987, KL<=0.0013"

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9c490cfb769bb4bce8cd979c9e96a84128be7bc8 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T103614Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2

20261004T103641Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72

! fail stability T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
Started2026-10-04 20:36 AEST
Finished2026-10-04 20:47 AEST (10.5 min)
Exit status0
Notesstability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 67 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 37.939 < floor 50.0

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Digestsha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max36,327 tokens
context.headroom_min168,259 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls67 calls {"per_min": 6.7, "requests": 46}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /v1/tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps37.9 tok/s
window 1: 39.97 tok/swindow 2: 43.13 tok/swindow 3: 42.36 tok/swindow 4: 46.42 tok/swindow 5: 43.39 tok/swindow 6: 42.38 tok/smin 39.97 · max 46.42 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 313, "min_at_generation_s": 87.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps44.1 tok/s {"generation_seconds": 432.859, "generated_tokens": 19093.9}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 216.5, "requests": 17, "tool_calls": 29, "completion_tokens": 6657, "occupied_max": 13921}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 383.5, "requests": 30, "tool_calls": 38, "completion_tokens": 12541, "occupied_max": 36327}] list
tasks_over_64k[] instance ids
completion_tokens19,198 tokens {"finish_reasons": {"tool_calls": 46}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 67 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 37.9 tok/s, whole run 44.1 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9c490cfb769bb4bce8cd979c9e96a84128be7bc8 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T103641Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2

20261004T104713Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72

✓ pass lab T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2
Started2026-10-04 20:47 AEST
Finished2026-10-04 20:52 AEST (5.3 min)
Exit status0
Notesserved=flashnext

Other runs of this config: load lab lab kit sweep TTFT stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T101015Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Digestsha256:f77c72f659567d28ffa2804806f69624e227530b4870a3d609f166014af1b451
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6.json sha256 1671d4e04532370bdb34ebf4958454f7b186984f5668d6c90538d416b0bca440
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max166,667 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps54.5 tok/s
lab_prefill_tps2,596 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9c490cfb769bb4bce8cd979c9e96a84128be7bc8 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T104713Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2

20261004T105250Z-sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2-load

✕ crash load T15 eligibility: none failed stage: load

Configsglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2
Started2026-10-04 20:52 AEST
Finished2026-10-04 20:55 AEST (3.0 min)
Exit status0
Notes_glue/offload_moe_method.py", line 86, in get_runtime store = HostExpertStore(model_dir, layers, num_experts, hidden, inter, bits, prefix=prefix) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/opt/trellis-serve/cuda/src/sglang_exl3/kernels/offload_store.py", line 62, in __init__ raise NotImplementedError("HostExpertStore: fast relayout implemented for K=3 experts only") NotImplementedError: HostExpertStore: fast relayout implemented for K=3 experts only [2026-10-04 10:55:46] Received sigquit from a child process. It usually means the child failed. [2026-10-04 10:55:46] kill_process_tree called: parent_pid=1, include_parent=True, pid=1 [2026-10-04 10:55:46] kill_process_tree called: parent_pid=1, include_parent=False, pid=1

Launch

recorded by tools/run.py (artifacts/runs/20261004T105250Z-lab-sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4:/models/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Digestsha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb.json sha256 0d170921f6e46f05cad846a1404d158b6e85e4fffe18eeec9a2d7d70ec79a969
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@65c895314393431c09050b2e04e250836b3a6eb4 · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-2.05bpw_h4_ng4.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_secondsunavailable not produced (crash/load)
host_ram_drop_gbunavailable not produced (crash/load)
vram_ready_mibunavailable not produced (crash/load)

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitdf0d10f31ae46a27f371fda8e1f15e40dfa878f0 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T105250Z-sglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2-load.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2

20261004T105602Z-sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2-load

✕ crash load T15 eligibility: none failed stage: load

Configsglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2
Started2026-10-04 20:56 AEST
Finished2026-10-04 20:56 AEST (0.9 min)
Exit status0
Notes ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/opt/trellis-serve/cuda/src/sglang_exl3/offload/ngram_nvme.py", line 116, in __init__ t = parse_table(model_path) ^^^^^^^^^^^^^^^^^^^^^^^ File "/opt/trellis-serve/cuda/src/sglang_exl3/offload/ngram_nvme.py", line 94, in parse_table raise ValueError(f"{path}: shard_N.trellis tensors missing or not contiguous") ValueError: /models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6/ngram_embedding.safetensors: shard_N.trellis tensors missing or not contiguous [2026-10-04 10:56:54] Received sigquit from a child process. It usually means the child failed. [2026-10-04 10:56:54] kill_process_tree called: parent_pid=1, include_parent=True, pid=1 [2026-10-04 10:56:54] kill_process_tree called: parent_pid=1, include_parent=False, pid=1

Launch

recorded by tools/run.py (artifacts/runs/20261004T105602Z-lab-sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6:/models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Digestsha256:ce47df25cee7ff235d8bfdf3e79e3828aa11bee3624db3ce121ebc79a8375dd3
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/1e72b250cb529c6d81bf11b88740d9a0b22f62de
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb.json sha256 ce41d62d49f95e441684dbb7afcbfd2cd54c0cdf49476dda70cda1eca5bc7928
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 8192 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@55a732e0c4c3d4614bc42b68493bb930d9b02c0a · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_secondsunavailable not produced (crash/load)
host_ram_drop_gbunavailable not produced (crash/load)
vram_ready_mibunavailable not produced (crash/load)

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitbd62e89c96f97db97945fd89be01017a19fea7c3 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T105602Z-sglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2-load.json

Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2

20261004T105700Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-load

✓ pass load T23 eligibility: none

Configexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
Started2026-10-04 20:57 AEST
Finished2026-10-04 21:02 AEST (5.0 min)
Exit statusNone
Notes–

Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T105700Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
Imageghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Digestsha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Engineexllamav3
Engine sourcehttps://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa
Launch filelaunches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17fa
MTP / speculativeoff
argv/opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
envGLM53_MODE=exact
GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw
GLM53_MODEL_DOWNLOAD=0
HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3@332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds300 s
host_ram_drop_gb120 GB
vram_ready_mib23,322 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitd36a5ce7e63d3abd393f2511b2a55930e67bf549 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T105700Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-load.json

Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2

20261004T110201Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-lab

! fail lab T23 eligibility: none failed stage: gates

Configexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
Started2026-10-04 21:02 AEST
Finished2026-10-04 21:12 AEST (11.0 min)
Exit status0
Notesserved=glm-5.3-flash

Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T105700Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
Imageghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Digestsha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Engineexllamav3
Engine sourcehttps://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa
Launch filelaunches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17fa
MTP / speculativeoff
argv/opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
envGLM53_MODE=exact
GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw
GLM53_MODEL_DOWNLOAD=0
HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3@332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps7.40 tok/s
lab_prefill_tps739 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitd36a5ce7e63d3abd393f2511b2a55930e67bf549 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T110201Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-lab.json

Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2

20261004T111300Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-ttft

! fail TTFT T23 eligibility: none failed stage: invalid measurement

Configexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
Started2026-10-04 21:13 AEST
Finished2026-10-04 21:17 AEST (4.2 min)
Exit status0
NotesINVALIDATED: TTFT measured to an empty role chunk the exllamav3 server sends before prefill (ttft.py v1 bug); re-measured with v2. headline length 8192; prompts carry random nonces

Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T105700Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
Imageghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Digestsha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Engineexllamav3
Engine sourcehttps://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa
Launch filelaunches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17fa
MTP / speculativeoff
argv/opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
envGLM53_MODE=exact
GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw
GLM53_MODEL_DOWNLOAD=0
HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3@332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tpsunavailable invalidated: TTFT measured to an empty role chunk the exllamav3 server sends before prefill (ttft.py v1 bug); re-measured with v2
ttft_prefill_32k_tpsunavailable invalidated: TTFT measured to an empty role chunk the exllamav3 server sends before prefill (ttft.py v1 bug); re-measured with v2
ttft_prefill_tpsunavailable invalidated: TTFT measured to an empty role chunk the exllamav3 server sends before prefill (ttft.py v1 bug); re-measured with v2

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitd36a5ce7e63d3abd393f2511b2a55930e67bf549 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T111300Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-ttft.json

Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2

20261004T111712Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-stability

! fail stability T23 eligibility: none

Configexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
Started2026-10-04 21:17 AEST
Finished2026-10-04 21:27 AEST (10.5 min)
Exit status0
Notesstability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 21 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 3.356 < floor 15.0

Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T105700Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
Imageghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Digestsha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Engineexllamav3
Engine sourcehttps://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa
Launch filelaunches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17fa
MTP / speculativeoff
argv/opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
envGLM53_MODE=exact
GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw
GLM53_MODEL_DOWNLOAD=0
HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3@332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json

Metrics

context.configured131,072 tokens
context.occupied_max4,589 tokens
context.headroom_min126,405 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls21 calls {"per_min": 2.1, "requests": 14}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps3.36 tok/s
window 1: 6.08 tok/swindow 2: 5.96 tok/smin 5.96 · max 6.08 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 99, "min_at_generation_s": 101.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps6.56 tok/s {"generation_seconds": 218.025, "generated_tokens": 1431.1}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 385.2, "requests": 11, "tool_calls": 16, "completion_tokens": 1605, "occupied_max": 4589}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 214.8, "requests": 4, "tool_calls": 5, "completion_tokens": 179, "occupied_max": 2222}] list
tasks_over_64k[] instance ids
completion_tokens1,784 tokens {"finish_reasons": {"tool_calls": 14}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 21 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 3.36 tok/s, whole run 6.56 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitd36a5ce7e63d3abd393f2511b2a55930e67bf549 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T111712Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-stability.json

Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2

20261004T112744Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-lab

! fail lab T23 eligibility: none failed stage: gates

Configexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
Started2026-10-04 21:27 AEST
Finished2026-10-04 22:05 AEST (37.9 min)
Exit status0
Notesserved=glm-5.3-flash

Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T105700Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
Imageghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Digestsha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Engineexllamav3
Engine sourcehttps://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa
Launch filelaunches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17fa
MTP / speculativeoff
argv/opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
envGLM53_MODE=exact
GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw
GLM53_MODEL_DOWNLOAD=0
HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3@332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json

Metrics

context.configured131,072 tokens
context.occupied_max93,437 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": false}
lab_decode_c1_tps7.50 tok/s
lab_prefill_tps720 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitd36a5ce7e63d3abd393f2511b2a55930e67bf549 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T112744Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-lab.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2

20261004T120556Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass load T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Started2026-10-04 22:05 AEST
Finished2026-10-04 22:10 AEST (4.3 min)
Exit statusNone
Notes–

Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T120556Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds260 s
host_ram_drop_gb65.9 GB
vram_ready_mib20,822 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitdcbaa681afe8032d6e59244d27f56918b1e7bd20 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T120556Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2

20261004T121018Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass lab T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Started2026-10-04 22:10 AEST
Finished2026-10-04 22:15 AEST (4.7 min)
Exit status0
Notesserved=flashnext

Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T120556Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max166,667 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps55.8 tok/s
lab_prefill_tps2,419 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitdcbaa681afe8032d6e59244d27f56918b1e7bd20 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T121018Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2

20261004T121501Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass kit sweep T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Started2026-10-04 22:15 AEST
Finished2026-10-04 22:30 AEST (15.7 min)
Exit status0
Notessweep status: DONE

Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T120556Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
prefill_8k_tps2,281 tok/s {"min": 2261.8, "max": 2289.8, "n": 3}
prefill_16k_tps2,237 tok/s {"min": 2234.0, "max": 2244.7, "n": 3}
prefill_32k_tps2,678 tok/s {"min": 2670.7, "max": 2679.3, "n": 3}
prefill_64k_tps2,649 tok/s {"min": 2643.2, "max": 2654.6, "n": 3}
decode_c1_tps47.4 tok/s {"min": 46.48, "max": 48.27, "power_w": null}
decode_c2_tps53.1 tok/s {"min": 51.34, "max": 54.86, "power_w": null}
decode_c3_tps60.6 tok/s {"min": 57.47, "max": 63.7, "power_w": null}
decode_c4_tps53.7 tok/s {"min": 52.62, "max": 54.77, "power_w": null}
decode_c1_32k_tps48.5 tok/s {"min": 48.52, "max": 48.57, "power_w": null}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitdcbaa681afe8032d6e59244d27f56918b1e7bd20 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T121501Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2

20261004T123041Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass TTFT T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Started2026-10-04 22:30 AEST
Finished2026-10-04 22:32 AEST (1.4 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces; TTFT = first chunk with generated output (v2)

Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T120556Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps2,009 tok/s {"prompt_tokens": [13741, 13755, 13738], "ttft_s": [6.864, 6.84, 6.838]}
ttft_prefill_32k_tps2,681 tok/s {"prompt_tokens": [54658, 54711, 54684], "ttft_s": [20.392, 20.395, 20.396]}
ttft_prefill_tps2,009 tok/s {"prompt_tokens": [13741, 13755, 13738], "ttft_s": [6.864, 6.84, 6.838]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitdcbaa681afe8032d6e59244d27f56918b1e7bd20 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T123041Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2

20261004T123204Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass quality panel T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Started2026-10-04 22:32 AEST
Finished2026-10-04 22:32 AEST (0.4 min)
Exit status0
Notespanel=qwen3.8-flash-next-exl3-ref-panel.json rc=0

Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T120556Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
top1_agreement0.98967 fraction
mean_kl0.00097327 nats
in_bandyes "top1>=0.987, KL<=0.0013"

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitdcbaa681afe8032d6e59244d27f56918b1e7bd20 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T123204Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2

20261004T123230Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✕ crash stability T13 eligibility: none failed stage: measure

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Started2026-10-04 22:32 AEST
Finished2026-10-04 22:42 AEST (10.1 min)
Exit status0
Notesstability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); decode floor not verifiable (too little generation)

Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T120556Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no successful request
context.headroom_minunavailable configured context unknown (set CONFIGURED_CTX)
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls0 calls {"per_min": 0.0, "requests": 0}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": null, "saved": "artifacts: proxy/repetition/"}
decode_window_min_tpsunavailable only 0.0 s of generation; need > 120 s for one window
decode_whole_run_tpsunavailable no generation recorded
tasks_attempted3 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved0 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomyes {"upstream_connection_failures": 77, "upstream_http_errors": 1, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Error:BadGatewayError: litellm.BadGatewayError: BadGatewayError: OpenAIException - Cannot connect to host 127.0.0.1:18080 ssl:default [Connect call failed ('127.0.0.1', 18080)]", "submitted": false, "resolved": false, "seconds": 267.3}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "Error:BadGatewayError: litellm.BadGatewayError: BadGatewayError: OpenAIException - Cannot connect to host 127.0.0.1:18080 ssl:default [Connect call failed ('127.0.0.1', 18080)]", "submitted": false, "resolved": false, "seconds": 261.8}, {"attempt": 3, "round": 1, "instance_id": "psf__requests-2317", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 71.0}] list
tasks_over_64k[] instance ids
completion_tokens0 tokens {"finish_reasons": {}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 0 tool calls, 0 parse failures, 0 repetition hits, tasks solved 0/3, decode min window – tok/s, whole run – tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitdcbaa681afe8032d6e59244d27f56918b1e7bd20 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T123230Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2

20261004T124252Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass load T14a eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2
Started2026-10-04 22:42 AEST
Finished2026-10-04 22:47 AEST (4.3 min)
Exit statusNone
Notes–

Other runs of this config: load lab ✕ TTFT ✕

Launch

recorded by tools/run.py (artifacts/runs/20261004T124252Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=6.0 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6.json sha256 90d5526c99bf2c3b87f25e51276eed60130a38e3a34bc2991695048a8bb33eb9
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=6.0
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds257 s
host_ram_drop_gb65.3 GB
vram_ready_mib21,706 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitdb208909a5ae27d314ff09d53d7cf803d67fda9c dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T124252Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2

20261004T124710Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✕ crash lab T14a eligibility: none failed stage: harness

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2
Started2026-10-04 22:47 AEST
Finished2026-10-04 22:48 AEST (1.7 min)
Exit status1
Notes–

Other runs of this config: load lab ✕ TTFT ✕

Launch

recorded by tools/run.py (artifacts/runs/20261004T124252Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=6.0 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6.json sha256 90d5526c99bf2c3b87f25e51276eed60130a38e3a34bc2991695048a8bb33eb9
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=6.0
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
gates_passedunavailable not produced (crash/harness)
lab_decode_c1_tpsunavailable not produced (crash/harness)
lab_prefill_tpsunavailable not produced (crash/harness)
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitdb208909a5ae27d314ff09d53d7cf803d67fda9c dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T124710Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2

20261004T124852Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✕ crash TTFT T14a eligibility: none failed stage: measure

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2
Started2026-10-04 22:48 AEST
Finished2026-10-04 22:49 AEST (0.3 min)
Exit status1
Notes–

Other runs of this config: load lab ✕ TTFT ✕

Launch

recorded by tools/run.py (artifacts/runs/20261004T124252Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=6.0 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6.json sha256 90d5526c99bf2c3b87f25e51276eed60130a38e3a34bc2991695048a8bb33eb9
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=6.0
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_tpsunavailable not produced (crash/measure)

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commitdb208909a5ae27d314ff09d53d7cf803d67fda9c dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T124852Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2

20261004T124925Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-load

✓ pass load T23 eligibility: none

Configexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
Started2026-10-04 22:49 AEST
Finished2026-10-04 22:54 AEST (4.6 min)
Exit statusNone
Notes–

Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T124925Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
Imageghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Digestsha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Engineexllamav3
Engine sourcehttps://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa
Launch filelaunches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17fa
MTP / speculativeoff
argv/opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
envGLM53_MODE=exact
GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw
GLM53_MODEL_DOWNLOAD=0
HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3@332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds276 s
host_ram_drop_gb121 GB
vram_ready_mib23,322 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commite2424ec92249689799954b79c057d915d7f402f1 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T124925Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-load.json

Runs / GLM-5.3-Flash / exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2

20261004T125401Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-ttft

✓ pass TTFT T23 eligibility: none

Configexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2
Started2026-10-04 22:54 AEST
Finished2026-10-04 22:58 AEST (4.2 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces; TTFT = first chunk with generated output (v2)

Other runs of this config: load load lab ! lab ! TTFT ! TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T124925Z-lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2/docker-run.txt)
docker run -d --name lab-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30000 --shm-size 16g -e GLM53_MODE=exact -e GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -e GLM53_MODEL_DOWNLOAD=0 -e HF_HUB_OFFLINE=1 -v ~/models/local-ai-exp/turboderp-GLM-5.3-Flash-exl3-3.05bpw:/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw:ro --entrypoint /opt/glm53/docker/entrypoint.sh ghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
Imageghcr.io/0xsero/glm53-flash-offload@sha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Digestsha256:bb633b0bcb85573ad40b6c408af5e1036ae062c593e42479f4e4d2ab521e551d
Engineexllamav3
Engine sourcehttps://github.com/0xSero/glm53-flash-offload/tree/fe97bcf846347d035859a0610cebbb707fa36caa
Launch filelaunches/exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x.json sha256 b679beb9863b299d0b6e44b1269b954e5292fb8039caa75a0e4c0d34e6cb17fa
MTP / speculativeoff
argv/opt/glm53/docker/entrypoint.sh -m /models/turboderp-GLM-5.3-Flash-exl3-3.05bpw -cs 131072 --max-batch-size 8 -chunk_size 8192 -ambs 4 --host 0.0.0.0 --port 30000 --served-name glm-5.3-flash
envGLM53_MODE=exact
GLM53_MODEL_DIR=/models/turboderp-GLM-5.3-Flash-exl3-3.05bpw
GLM53_MODEL_DOWNLOAD=0
HF_HUB_OFFLINE=1
turboderp/GLM-5.3-Flash-exl3@332ab457b709b7ba30dd9a448be5de03b80a7ac9 · manifest models/turboderp-GLM-5.3-Flash-exl3-3.05bpw.json

Metrics

context.configured131,072 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps483 tok/s {"prompt_tokens": [10753, 10793, 10717], "ttft_s": [19.173, 22.332, 25.068]}
ttft_prefill_32k_tps688 tok/s {"prompt_tokens": [42865, 42692, 42811], "ttft_s": [62.271, 62.501, 60.752]}
ttft_prefill_tps483 tok/s {"prompt_tokens": [10753, 10793, 10717], "ttft_s": [19.173, 22.332, 25.068]}

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commite2424ec92249689799954b79c057d915d7f402f1 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/glm-5.3-flash/20261004T125401Z-exllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2-ttft.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2

20261004T125827Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass load T14a eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Started2026-10-04 22:58 AEST
Finished2026-10-04 23:02 AEST (4.4 min)
Exit statusNone
Notes–

Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T125827Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds264 s
host_ram_drop_gb65.7 GB
vram_ready_mib20,844 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9229192eace16fc1782f479846e0194aea593ba9 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T125827Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2

20261004T130251Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

! fail stability T14a eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Started2026-10-04 23:02 AEST
Finished2026-10-04 23:13 AEST (10.5 min)
Exit status0
Notesstability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 71 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 39.676 < floor 50.0

Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T125827Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max25,309 tokens
context.headroom_min178,942 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls71 calls {"per_min": 7.1, "requests": 47}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /v1/tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps39.7 tok/s
window 1: 44.31 tok/swindow 2: 43.52 tok/swindow 3: 43.73 tok/swindow 4: 44.87 tok/swindow 5: 41.92 tok/smin 41.92 · max 44.87 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 254, "min_at_generation_s": 267.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps43.3 tok/s {"generation_seconds": 373.189, "generated_tokens": 16176.9}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 245.6, "requests": 21, "tool_calls": 31, "completion_tokens": 5704, "occupied_max": 13986}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 354.4, "requests": 27, "tool_calls": 40, "completion_tokens": 10614, "occupied_max": 25309}] list
tasks_over_64k[] instance ids
completion_tokens16,318 tokens {"finish_reasons": {"tool_calls": 47}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 71 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 39.7 tok/s, whole run 43.3 tok/s.

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9229192eace16fc1782f479846e0194aea593ba9 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T130251Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2

20261004T131323Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass lab T14a eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Started2026-10-04 23:13 AEST
Finished2026-10-04 23:22 AEST (8.9 min)
Exit status0
Notesserved=flashnext

Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T125827Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max166,667 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps57.4 tok/s
lab_prefill_tps2,537 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit9229192eace16fc1782f479846e0194aea593ba9 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T131323Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2

20261004T132232Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass load T14a eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2
Started2026-10-04 23:22 AEST
Finished2026-10-04 23:26 AEST (4.2 min)
Exit statusNone
Notes–

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T132232Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1.json sha256 e76072403dd6ef64022ca8b6ccb17b7e7d25f20aa7c17b8287c7534158eee184
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds254 s
host_ram_drop_gb64.4 GB
vram_ready_mib20,744 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit97294db30283c264f2c58088bb36bfb1619ba3de dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T132232Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2

20261004T132647Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass lab T14a eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2
Started2026-10-04 23:26 AEST
Finished2026-10-04 23:31 AEST (5.0 min)
Exit status0
Notesserved=flashnext

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T132232Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1.json sha256 e76072403dd6ef64022ca8b6ccb17b7e7d25f20aa7c17b8287c7534158eee184
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max166,667 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps55.6 tok/s
lab_prefill_tps2,430 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit97294db30283c264f2c58088bb36bfb1619ba3de dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T132647Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2

20261004T133146Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass TTFT T14a eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2
Started2026-10-04 23:31 AEST
Finished2026-10-04 23:33 AEST (1.4 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces; TTFT = first chunk with generated output (v2)

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T132232Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1.json sha256 e76072403dd6ef64022ca8b6ccb17b7e7d25f20aa7c17b8287c7534158eee184
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps2,014 tok/s {"prompt_tokens": [13689, 13759, 13721], "ttft_s": [6.848, 6.829, 6.812]}
ttft_prefill_32k_tps2,692 tok/s {"prompt_tokens": [54692, 54821, 54730], "ttft_s": [20.316, 20.367, 20.356]}
ttft_prefill_tps2,014 tok/s {"prompt_tokens": [13689, 13759, 13721], "ttft_s": [6.848, 6.829, 6.812]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit97294db30283c264f2c58088bb36bfb1619ba3de dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T133146Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2

20261004T133309Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

! fail stability T14a eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2
Started2026-10-04 23:33 AEST
Finished2026-10-04 23:43 AEST (10.5 min)
Exit status0
Notesstability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 57 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 40.08 < floor 50.0

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T132232Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1.json sha256 e76072403dd6ef64022ca8b6ccb17b7e7d25f20aa7c17b8287c7534158eee184
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max29,027 tokens
context.headroom_min174,880 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls57 calls {"per_min": 5.7, "requests": 42}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /v1/tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps40.1 tok/s
window 1: 42.01 tok/swindow 2: 46.09 tok/swindow 3: 40.09 tok/swindow 4: 43.71 tok/smin 40.09 · max 46.09 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 193, "min_at_generation_s": 181.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps43.1 tok/s {"generation_seconds": 312.161, "generated_tokens": 13447.2}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 124.1, "requests": 11, "tool_calls": 18, "completion_tokens": 3257, "occupied_max": 8494}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 475.9, "requests": 32, "tool_calls": 39, "completion_tokens": 10303, "occupied_max": 29027}] list
tasks_over_64k[] instance ids
completion_tokens13,560 tokens {"finish_reasons": {"tool_calls": 42}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 57 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 40.1 tok/s, whole run 43.1 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit97294db30283c264f2c58088bb36bfb1619ba3de dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T133309Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2

20261004T134340Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass lab T14a eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2
Started2026-10-04 23:43 AEST
Finished2026-10-04 23:48 AEST (4.6 min)
Exit status0
Notesserved=flashnext

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T132232Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1.json sha256 e76072403dd6ef64022ca8b6ccb17b7e7d25f20aa7c17b8287c7534158eee184
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 1 --cuda-graph-max-bs-decode 1 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max166,667 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps54.6 tok/s
lab_prefill_tps2,580 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit97294db30283c264f2c58088bb36bfb1619ba3de dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T134340Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2

20261004T134832Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass load T14b eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2
Started2026-10-04 23:48 AEST
Finished2026-10-04 23:52 AEST (3.6 min)
Exit statusNone
Notes–

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T134832Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=pinned -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned.json sha256 06f6cc13b0c25dfb2d50a8b515c6240bf1418dfc8b07facf231d3e0847358cb1
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=pinned
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds215 s
host_ram_drop_gb88.2 GB
vram_ready_mib20,740 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit1b7b92a2ae73d921b9b6ab2ff7cb2059734faf95 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T134832Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2

20261004T135208Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass lab T14b eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2
Started2026-10-04 23:52 AEST
Finished2026-10-04 23:55 AEST (3.7 min)
Exit status0
Notesserved=flashnext

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T134832Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=pinned -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned.json sha256 06f6cc13b0c25dfb2d50a8b515c6240bf1418dfc8b07facf231d3e0847358cb1
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=pinned
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max166,667 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps53.7 tok/s
lab_prefill_tps2,440 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit1b7b92a2ae73d921b9b6ab2ff7cb2059734faf95 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T135208Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2

20261004T135547Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass TTFT T14b eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2
Started2026-10-04 23:55 AEST
Finished2026-10-04 23:57 AEST (1.4 min)
Exit status0
Notesheadline length 8192; prompts carry random nonces; TTFT = first chunk with generated output (v2)

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T134832Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=pinned -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned.json sha256 06f6cc13b0c25dfb2d50a8b515c6240bf1418dfc8b07facf231d3e0847358cb1
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=pinned
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
ttft_prefill_8k_tps2,022 tok/s {"prompt_tokens": [13765, 13739, 13758], "ttft_s": [6.807, 6.819, 6.799]}
ttft_prefill_32k_tps2,700 tok/s {"prompt_tokens": [54781, 54752, 54783], "ttft_s": [20.259, 20.275, 20.301]}
ttft_prefill_tps2,022 tok/s {"prompt_tokens": [13765, 13739, 13758], "ttft_s": [6.807, 6.819, 6.799]}

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit1b7b92a2ae73d921b9b6ab2ff7cb2059734faf95 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T135547Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2

20261004T135709Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

! fail stability T14b eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2
Started2026-10-04 23:57 AEST
Finished2026-10-05 00:07 AEST (10.5 min)
Exit status0
Notesstability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 72 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 40.915 < floor 50.0

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T134832Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=pinned -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned.json sha256 06f6cc13b0c25dfb2d50a8b515c6240bf1418dfc8b07facf231d3e0847358cb1
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=pinned
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max39,281 tokens
context.headroom_min165,395 tokens
duration_s600 s {"cap_s": 600.0, "driver_capped": true}
tool_calls72 calls {"per_min": 7.2, "requests": 50}
parse_failures0 responses {"wire": {}, "harness_format_errors": {}, "union_rule": "a response counts once even if both the wire check and the harness flag it", "counted": "malformed tool calls only (decision 0003); truncations and other format errors are separate metrics"}
output_truncations0 responses {"rule": "cut off at max_tokens with no tool call"}
agent_format_errors0 responses {"rule": "agent protocol errors other than malformed tool calls"}
repetition_hits0 hits {"ngram64x3": 0, "same_call_x4": 0, "ngram64x3_cross_field_not_counted": 0, "tokenizer": "server /v1/tokenize", "saved": "artifacts: proxy/repetition/"}
decode_window_min_tps40.9 tok/s
window 1: 46.31 tok/swindow 2: 42.86 tok/swindow 3: 45.63 tok/swindow 4: 44.46 tok/swindow 5: 47.12 tok/smin 42.86 · max 47.12 tok/s
{"window_s": 60, "skip_s": 60, "rolling_step_s": 1, "n_rolling_windows": 289, "min_at_generation_s": 108.0, "windows_tps_note": "consecutive non-overlapping 60 s windows; the value is the minimum over rolling windows (1 s step)"}
decode_whole_run_tps44.6 tok/s {"generation_seconds": 408.606, "generated_tokens": 18207.3}
tasks_attempted2 attempts {"unfinished_at_cap": 1, "task_ids": ["django__django-11099", "sympy__sympy-21612", "psf__requests-2317", "matplotlib__matplotlib-23299", "scikit-learn__scikit-learn-13439", "astropy__astropy-12907", "pytest-dev__pytest-7373", "sympy__sympy-20590"], "task_set_sha256": "8e7487ca7d059df6359fc7d9b86c7021256a8a81b8af049616750f8f83400324"}
tasks_solved1 attempts {"eval_errors": 0, "swebench": "5.0.2"}
crash_or_oomno {"upstream_connection_failures": 0, "upstream_http_errors": 0, "source": "proxy-observed; the run wrapper adds container death/OOM"}
gates_afterskipped filled by qualify step: the lab run that follows this run (tools/qualify.py)
agent_attempts[{"attempt": 1, "round": 1, "instance_id": "django__django-11099", "exit_status": "Submitted", "submitted": true, "resolved": true, "seconds": 137.5, "requests": 16, "tool_calls": 22, "completion_tokens": 3550, "occupied_max": 10227}, {"attempt": 2, "round": 1, "instance_id": "sympy__sympy-21612", "exit_status": "WallClockCap", "submitted": false, "resolved": false, "seconds": 462.8, "requests": 34, "tool_calls": 50, "completion_tokens": 14781, "occupied_max": 39281}] list
tasks_over_64k[] instance ids
completion_tokens18,331 tokens {"finish_reasons": {"tool_calls": 50}}
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Agentic summary: 600 s, 72 tool calls, 0 parse failures, 0 repetition hits, tasks solved 1/2, decode min window 40.9 tok/s, whole run 44.6 tok/s.

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit1b7b92a2ae73d921b9b6ab2ff7cb2059734faf95 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T135709Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2

20261004T140742Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass lab T14b eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2
Started2026-10-05 00:07 AEST
Finished2026-10-05 00:11 AEST (3.7 min)
Exit status0
Notesserved=flashnext

Other runs of this config: load lab lab TTFT stability !

Launch

recorded by tools/run.py (artifacts/runs/20261004T134832Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=pinned -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned.json sha256 06f6cc13b0c25dfb2d50a8b515c6240bf1418dfc8b07facf231d3e0847358cb1
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=pinned
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_max166,667 tokens
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps51.3 tok/s
lab_prefill_tps2,596 tok/s "context-gate prompt tokens / total request seconds"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachewarm
Layoutcurrent · 1 card(s) used · display off
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8off
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit1b7b92a2ae73d921b9b6ab2ff7cb2059734faf95 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T140742Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2

20261004T141422Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass load T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Started2026-10-05 00:14 AEST
Finished2026-10-05 00:18 AEST (4.3 min)
Exit statusNone
Notes–

Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T141422Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
load_seconds257 s
host_ram_drop_gb64.5 GB
vram_ready_mib20,818 MiB

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit742e81596c31f25802bdf11347fc6d9ab1bb8265 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T141422Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Runs / Qwen3.8-Flash-Next / sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2

20261004T141839Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe

✓ pass lab T13 eligibility: none

Configsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2
Started2026-10-05 00:18 AEST
Finished2026-10-05 00:25 AEST (7.0 min)
Exit status0
Notesofficial lab.py try --on endpoint; run file rtx-3090-24gb.qwen3.8-flash-next.sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb.200k.20261004T142537.json; recipe written only on pass

Other runs of this config: load load load lab lab lab kit sweep TTFT stability ✕ stability ! quality panel

Launch

recorded by tools/run.py (artifacts/runs/20261004T141422Z-lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090/docker-run.txt)
docker run -d --name lab-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090 --security-opt no-new-privileges --gpus device=GPU-76d3c6af-7f18-69d2-3145-899103de1722 -p 127.0.0.1:18080:30100 --shm-size 64g -e HF_HUB_OFFLINE=1 -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 -e SGLANG_EXL3_MOE_OFFLOAD=gpu_cache -e EXL3_MOE_CPU_THREADS=24 -e SGLANG_EXL3_EXPERT_CACHE_GB=auto -e SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26 -e SGLANG_EXL3_KV_BITS=5 -e SGLANG_EXL3_EMBED_HOST=1 -e SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4 -e SGLANG_EXL3_OFFLOAD_FUSED=1 -e SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1 -e SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1 -e SGLANG_EXL3_NGRAM_TIER=nvme -e SGLANG_EXL3_VISION=1 -e SGLANG_EXL3_MM_FAST_CPU=1 -e SGLANG_EXL3_VIT_SDPA=1 -e SGLANG_EXL3_VIT_MLP_CHUNK=8192 -v ~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5:ro --entrypoint /opt/entrypoint.sh ghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163 python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
Imageghcr.io/0xsero/sglang-exl3-flashnext@sha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Digestsha256:a21efea6fe15ccc6a8774b1ca2a4c4104388cdcca8f1463645c02c0f8aba0163
Enginesglang
Engine sourcehttps://github.com/0xSero/trellis-serve/tree/b1e74afcf98203dc5326349623b556ba2df4b16a
Launch filelaunches/sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6.json sha256 cbcfae235adc9e4c7988a047150225a710b8e9ed2b2370701c5d2a9820c28a01
MTP / speculativeoff
argv/opt/entrypoint.sh python3 -m sglang.launch_server --model-path /models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5 --quantization exl3 --trust-remote-code --host 0.0.0.0 --port 30100 --served-model-name flashnext --disable-shared-experts-fusion --kv-cache-dtype fp8_e4m3 --context-length 204800 --mem-fraction-static 0.88 --chunked-prefill-size 12288 --max-running-requests 4 --cuda-graph-max-bs-decode 8 --cuda-graph-backend-prefill disabled --max-mamba-cache-size 16 --max-total-tokens 210000 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
envHF_HUB_OFFLINE=1
CUDA_DEVICE_ORDER=PCI_BUS_ID
SGLANG_EXL3_MODEL_PATH=/models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5
SGLANG_EXL3_MOE_OFFLOAD=gpu_cache
EXL3_MOE_CPU_THREADS=24
SGLANG_EXL3_EXPERT_CACHE_GB=auto
SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB=8.26
SGLANG_EXL3_KV_BITS=5
SGLANG_EXL3_EMBED_HOST=1
SGLANG_EXL3_OFFLOAD_STAGING_PARTS=4
SGLANG_EXL3_OFFLOAD_FUSED=1
SGLANG_EXL3_OFFLOAD_COMPACT_PARTS=1
SGLANG_EXL3_MOE_PREFILL_FP16_ACC=1
SGLANG_EXL3_NGRAM_TIER=nvme
SGLANG_EXL3_VISION=1
SGLANG_EXL3_MM_FAST_CPU=1
SGLANG_EXL3_VIT_SDPA=1
SGLANG_EXL3_VIT_MLP_CHUNK=8192
turboderp/Qwen3.8-Flash-Next-exl3@69e33439ae950f17bcbe95c98f117d80f759ab6d · manifest models/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5.json

Metrics

context.configured204,800 tokens
context.occupied_maxunavailable no occupancy measured in this run
context.headroom_minunavailable no occupancy measured in this run
gates_passed["load", "chat", "reasoning", "tools", "context", "speed"] gates {"load": true, "chat": true, "reasoning": true, "tools": true, "context": true, "speed": true}
lab_decode_c1_tps55.8 tok/s
lab_prefill_tps2,419 tok/s "official lab.py try"
mtp_accepted_tpsskipped speculative decoding off in this launch
mtp_acceptance_rateskipped speculative decoding off in this launch

Cache, layout and host

Cachecold · reset: container, prefix_kv, expert_cache, ngram_row_cache, page_cache
Layoutcurrent · 1 card(s) used · display on
hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
cpus_offline2,18
cpu_boostFalse
cpu_max_mhz3401
machine_checks_this_boot6
#UsedNameUUIDBusPCIeDisplay
0NVIDIA GeForce RTX 3090GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
1NVIDIA GeForce RTX 3090GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
2●NVIDIA GeForce RTX 3090GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

Provenance

bench commit742e81596c31f25802bdf11347fc6d9ab1bb8265 dirty tree
pins.json sha256858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78 (matches current pins.json)
registryd21258dd744e7c78be28177c90060c6af8e10b7e
kitef883d269e50ecbea290f919f095e2f3ca633b42
mini-swe-agent04d809ceab9df28f9adaed044884180159172930

Artifacts and raw output

Raw record: results/qwen3.8-flash-next/20261004T141839Z-sglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efe.json

Hardware & methodology

Host-to-device bandwidth

RunOutcomeLayoutGPUBusPCIeH2DD2HNotes
2026-10-03 06:53 AEST✓ passcurrentGPU2 GPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x813.5 GB/s13.2 GB/s1 GiB pinned transfers; full table in detail
2026-10-03 06:53 AEST✓ passcurrentGPU0 GPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x86.03 GB/s6.60 GB/s1 GiB pinned transfers; full table in detail
2026-10-04 19:11 AEST✓ passcurrentGPU1 GPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x813.5 GB/s13.2 GB/s1 GiB pinned transfers; full table in detail

NVMe random reads

RunOutcomeIOPSThroughputDetailNotes
2026-10-03 06:54 AEST✓ pass168,600 IOPS3.35 GB/s"16 KiB, 32 threads"~/models/local-ai-exp/turboderp-Qwen3.8-Flash-Next-exl3-3.05bpw_h5_ng5/ngram_embedding.safetensors

Host

hostnameomarchy-gpu
cpuAMD Ryzen 9 5950X 16-Core Processor
ram_total_gib125.7
ram_layoutDIMM_A1 32 GiB DDR4 3200 MT/s, DIMM_A2 32 GiB DDR4 3200 MT/s, DIMM_B1 32 GiB DDR4 3200 MT/s, DIMM_B2 32 GiB DDR4 3200 MT/s (AM4: 2 channels)
kernel7.2.5-3-omarchy
driver610.57.04
cuda13.3
cpus_online0-1,3-17,19-31
0-31 (changed between runs)
cpus_offline2,18
none (changed between runs)
machine_checks_this_boot0
6 (changed between runs)
cpu_boostFalse
cpu_max_mhz3401
LayoutGPU UUIDBusPCIeDisplay
currentGPU-816b43e4-8d65-dfd7-2234-8517b3dfbf2d00000000:04:00.0gen4 x8off
currentGPU-c67ac872-3371-a88c-aa92-d969daf6405e00000000:0B:00.0gen4 x8on
currentGPU-76d3c6af-7f18-69d2-3145-899103de172200000000:0C:00.0gen4 x8off

GPUs as seen by the records (PCIe width is the link width at record time).

Dropped configs and why

ModelConfigFailed runsWhy (record notes)
GLMexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2! fail lab / gates ! fail TTFT / invalid measurement ! fail stability ! fail lab / gatesserved=glm-5.3-flash
INVALIDATED: TTFT measured to an empty role chunk the exllamav3 server sends before prefill (ttft.py v1 bug); re-measured with v2. headline length 8192; prompts carry random nonces
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 21 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 3.356 < floor 15.0
served=glm-5.3-flash
GLMglm-llamacpp-iq2m-128k-1x-gpu2! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 24 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 8.91 < floor 15.0
served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
GLMik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 47 tool-call parse failure(s); decode window min 9.458 < floor 15.0
served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
GLMik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2! fail lab / gates ✕ crash stability / measureserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); 9 tool-call parse failure(s); decode floor not verifiable (too little generation)
GLMik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2! fail lab / gates ✕ crash stability / measureserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); 9 tool-call parse failure(s); decode floor not verifiable (too little generation)
GLMik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 68 tool-call parse failure(s); decode window min 13.255 < floor 15.0
served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
GLMik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012✕ crash load / loadche_init: CUDA2 KV buffer size = 570.81 MiB llama_init_from_model: KV self size = 748.00 MiB, c^KV (q8_0): 748.00 MiB, kv^T: not used llama_init_from_model: KV self indexer size = 704.00 MiB (f16) llama_init_from_model: CUDA_Host output buffer size = 0.59 MiB ggml_backend_cuda_buffer_type_alloc_buffer: allocating 3488.51 MiB on device 0: cudaMalloc failed: out of memory ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 3657965696 llama_init_from_model: failed to allocate compute buffers ~ggml_backend_cuda_context: have 0 graphs ~ggml_backend_cuda_context: have 0 graphs ~ggml_backend_cuda_context: have 0 graphs llama_init_from_gpt_params: error: failed to create context with model '/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf'
GLMik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012✕ crash load / load c^KV (q8_0): 68.00 MiB, kv^T: not used llama_init_from_model: KV self indexer size = 64.00 MiB (f16) llama_init_from_model: CUDA_Host output buffer size = 0.61 MiB ggml_backend_cuda_buffer_type_alloc_buffer: allocating 2754.25 MiB on device 0: cudaMalloc failed: out of memory ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 2888039424 llama_init_from_model: failed to allocate compute buffers ~ggml_backend_cuda_context: have 0 graphs ~ggml_backend_cuda_context: have 0 graphs ~ggml_backend_cuda_context: have 0 graphs common_speculative_init: failed to create MTP context srv init: failed to initialize recurrent speculative context ~ggml_backend_cuda_context: have 13 graphs ~ggml_backend_cuda_context: have 14 graphs ~ggml_backend_cuda_context: have 11 graphs
GLMik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012! fail lab / gates ✕ crash stability / measureserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); 9 tool-call parse failure(s); decode floor not verifiable (too little generation)
GLMik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012✕ crash load / loadche_init: CUDA2 KV buffer size = 570.81 MiB llama_init_from_model: KV self size = 816.00 MiB, c^KV (q8_0): 816.00 MiB, kv^T: not used llama_init_from_model: KV self indexer size = 768.00 MiB (f16) llama_init_from_model: CUDA_Host output buffer size = 0.61 MiB ggml_backend_cuda_buffer_type_alloc_buffer: allocating 3488.51 MiB on device 0: cudaMalloc failed: out of memory ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 3657965696 llama_init_from_model: failed to allocate compute buffers ~ggml_backend_cuda_context: have 0 graphs ~ggml_backend_cuda_context: have 0 graphs ~ggml_backend_cuda_context: have 0 graphs llama_init_from_gpt_params: error: failed to create context with model '/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf'
GLMllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 30 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 8.854 < floor 15.0
served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
GLMllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
GLMllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 27 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 1 repetition hit(s); decode window min 8.847 < floor 15.0
served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
GLMllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
GLMllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 24 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 11.035 < floor 15.0
served=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
GLMllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
GLMllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012! fail lab / gates ! fail stability ✕ crash lab / harnessserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 26 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 11.388 < floor 15.0
(no notes)
GLMllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ2_M/GLM-5.3-Flash-BF16-IQ2_M-00001-of-00004.gguf
GLMllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 25 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 7.646 < floor 15.0
served=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf
GLMllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 22 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 1 repetition hit(s); decode window min 9.855 < floor 15.0
served=/models/GLM-5.3-Flash-BF16-IQ3_XXS/GLM-5.3-Flash-BF16-IQ3_XXS-00001-of-00004.gguf
GLMllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 25 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 7.807 < floor 15.0
served=/models/GLM-5.3-Flash-BF16-IQ4_XS/GLM-5.3-Flash-BF16-IQ4_XS-00001-of-00005.gguf
GLMllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2! fail lab / gates ! fail stability ! fail lab / gatesserved=/models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf
stability run, 10 min cap, floor 15.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 19 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 8.8 < floor 15.0
served=/models/GLM-5.3-Flash-BF16-Q2_K/GLM-5.3-Flash-BF16-Q2_K-00001-of-00004.gguf
Flash-Nextfn-llamacpp-iq4xs-128k-1x-fit-gpu2! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 49 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: 1 repetition hit(s); decode window min 23.83 < floor 50.0
Flash-Nextfn-llamacpp-iq4xs-128k-1x-gpu2✕ crash load / loaduser, abort 0.18.994.031 E ggml_backend_cuda_buffer_type_alloc_buffer: allocating 65358.17 MiB on device 0: cudaMalloc failed: out of memory 0.18.994.037 E alloc_tensor_range: failed to allocate CUDA0 buffer of size 68533006336 0.19.139.695 E llama_model_load: error loading model: unable to allocate CUDA0 buffer 0.19.139.702 E llama_model_load_from_file_impl: failed to load model 0.19.139.708 E cmn common_init_: failed to load model '/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf' 0.19.139.713 E srv load_model: failed to load model, '/models/Qwen3.8-Flash-Next-IQ4_XS/Qwen3.8-Flash-Next-IQ4_XS-00001-of-00003.gguf' 0.19.139.715 I srv operator(): operator(): cleaning up before exit... 0.19.140.570 E srv llama_server: exiting due to model loading error
Flash-Nextfn-trellis-3.05-0xsero-main-x8-gpu2! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 63 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 38.077 < floor 50.0
Flash-Nextllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 70 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 38.626 < floor 50.0
Flash-Nextllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode floor not verifiable (too little generation); agent driver exited 1
Flash-Nextsglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2✕ crash load / load_glue/offload_moe_method.py", line 86, in get_runtime store = HostExpertStore(model_dir, layers, num_experts, hidden, inter, bits, prefix=prefix) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/opt/trellis-serve/cuda/src/sglang_exl3/kernels/offload_store.py", line 62, in __init__ raise NotImplementedError("HostExpertStore: fast relayout implemented for K=3 experts only") NotImplementedError: HostExpertStore: fast relayout implemented for K=3 experts only [2026-10-04 10:55:46] Received sigquit from a child process. It usually means the child failed. [2026-10-04 10:55:46] kill_process_tree called: parent_pid=1, include_parent=True, pid=1 [2026-10-04 10:55:46] kill_process_tree called: parent_pid=1, include_parent=False, pid=1
Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 57 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 40.08 < floor 50.0
Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2✕ crash stability / measure ! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 0 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: engine connection failures (crash/OOM); decode floor not verifiable (too little generation)
stability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 71 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 39.676 < floor 50.0
Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 72 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 40.915 < floor 50.0
Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2✕ crash lab / harness ✕ crash TTFT / measure(no notes)
(no notes)
Flash-Nextsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2! fail stabilitystability run, 10 min cap, floor 50.0 tok/s; mini-swe-agent 04d809ceab9d (benchmarks/swebench.yaml, native tool calls); sampling {"temperature": 0.7, "top_p": 0.95, "seed": 1234, "max_tokens": 8192} | repetition tokenizer: server /v1/tokenize | 67 tool calls in this run count toward the config's 100-call minimum (§5) | FAIL: decode window min 37.939 < floor 50.0
Flash-Nextsglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2✕ crash load / load ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/opt/trellis-serve/cuda/src/sglang_exl3/offload/ngram_nvme.py", line 116, in __init__ t = parse_table(model_path) ^^^^^^^^^^^^^^^^^^^^^^^ File "/opt/trellis-serve/cuda/src/sglang_exl3/offload/ngram_nvme.py", line 94, in parse_table raise ValueError(f"{path}: shard_N.trellis tensors missing or not contiguous") ValueError: /models/turboderp-Qwen3.8-Flash-Next-exl3-4.05bpw_h6_ng6/ngram_embedding.safetensors: shard_N.trellis tensors missing or not contiguous [2026-10-04 10:56:54] Received sigquit from a child process. It usually means the child failed. [2026-10-04 10:56:54] kill_process_tree called: parent_pid=1, include_parent=True, pid=1 [2026-10-04 10:56:54] kill_process_tree called: parent_pid=1, include_parent=False, pid=1

All failed and crashed runs (64)

StartedOutcomeModelKindConfigKey metricsStageTicket
2026-10-04 23:57 AEST! failFlash-Nextstabilitysglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-pinned-gpu2decode_window_min_tps 40.9 tok/s · tool_calls 72 callsT14b
2026-10-04 23:33 AEST! failFlash-Nextstabilitysglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-g1-gpu2decode_window_min_tps 40.1 tok/s · tool_calls 57 callsT14a
2026-10-04 23:02 AEST! failFlash-Nextstabilitysglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2decode_window_min_tps 39.7 tok/s · tool_calls 71 callsT14a
2026-10-04 22:48 AEST✕ crashFlash-NextTTFTsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2ttft_prefill_tps unavailablemeasureT14a
2026-10-04 22:47 AEST✕ crashFlash-Nextlabsglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-r6-gpu2lab_decode_c1_tps unavailable · lab_prefill_tps unavailableharnessT14a
2026-10-04 22:32 AEST✕ crashFlash-Nextstabilitysglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-a21efea6-gpu2decode_window_min_tps unavailable · tool_calls 0 callsmeasureT13
2026-10-04 21:27 AEST! failGLMlabexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2lab_decode_c1_tps 7.50 tok/s · lab_prefill_tps 720 tok/sgatesT23
2026-10-04 21:17 AEST! failGLMstabilityexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2decode_window_min_tps 3.36 tok/s · tool_calls 21 callsT23
2026-10-04 21:13 AEST! failGLMTTFTexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2ttft_prefill_tps unavailableinvalid measurementT23
2026-10-04 21:02 AEST! failGLMlabexllamav3-glm-5.3-flash-exl3-3.05bpw-exact-128k-1x-gpu2lab_decode_c1_tps 7.40 tok/s · lab_prefill_tps 739 tok/sgatesT23
2026-10-04 20:56 AEST✕ crashFlash-Nextloadsglang-qwen3.8-flash-next-exl3-4.05bpw-offload-200k-rtx-3090-24gb-gpu2load_seconds unavailableloadT15
2026-10-04 20:52 AEST✕ crashFlash-Nextloadsglang-qwen3.8-flash-next-exl3-2.05bpw-offload-200k-rtx-3090-24gb-gpu2load_seconds unavailableloadT15
2026-10-04 20:36 AEST! failFlash-Nextstabilitysglang-qwen3.8-flash-next-exl3-3.05bpw-offload-200k-rtx-3090-24gb-phase-d-f77c72f6-gpu2decode_window_min_tps 37.9 tok/s · tool_calls 67 callsT13
2026-10-04 19:38 AEST✕ crashGLMstabilityik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012decode_window_min_tps unavailable · tool_calls 0 callsmeasureT22c
2026-10-04 19:12 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm7168-gpu012lab_decode_c1_tps 10.2 tok/s · lab_prefill_tps 166 tok/sgatesT22c
2026-10-04 17:42 AEST✕ crashGLMloadik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-fm4096-gpu012load_seconds unavailableloadT22c
2026-10-04 17:31 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012lab_decode_c1_tps 13.4 tok/s · lab_prefill_tps 189 tok/sgatesT22c
2026-10-04 17:21 AEST! failGLMstabilityik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012decode_window_min_tps 13.3 tok/s · tool_calls 0 callsT22c
2026-10-04 16:58 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-fm4096-gpu012lab_decode_c1_tps 14.0 tok/s · lab_prefill_tps 185 tok/sgatesT22c
2026-10-04 16:30 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub4096-fitt2048-g210-gpu012lab_decode_c1_tps 11.2 tok/s · lab_prefill_tps 258 tok/sgatesT21
2026-10-04 16:14 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub8192-gpu2lab_decode_c1_tps 9.10 tok/s · lab_prefill_tps 370 tok/sgatesT20b
2026-10-04 15:49 AEST! failFlash-Nextstabilityllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-fitt2048-g210-gpu012decode_window_min_tps 38.6 tok/s · tool_calls 70 callsT11
2026-10-04 15:19 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012lab_decode_c1_tps 8.30 tok/s · lab_prefill_tps 156 tok/sgatesT21
2026-10-04 15:08 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012decode_window_min_tps 7.81 tok/s · tool_calls 25 callsT21
2026-10-04 14:26 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq4-xs-128k-3x-ub2048-fitt2048-g210-gpu012lab_decode_c1_tps 8.00 tok/s · lab_prefill_tps 139 tok/sgatesT21
2026-10-04 14:24 AEST✕ crashGLMloadik_llama-glm-5.3-flash-iq2-m-128k-3x-mtp-g210-gpu012load_seconds unavailableloadT22c
2026-10-04 14:23 AEST✕ crashGLMloadik_llama-glm-5.3-flash-iq2-m-128k-3x-g210-gpu012load_seconds unavailableloadT22c
2026-10-04 14:14 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2lab_decode_c1_tps 9.60 tok/s · lab_prefill_tps 288 tok/sgatesT20b
2026-10-04 14:04 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2decode_window_min_tps 8.85 tok/s · tool_calls 27 callsT20b
2026-10-04 13:42 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub4096-gpu2lab_decode_c1_tps 9.60 tok/s · lab_prefill_tps 288 tok/sgatesT20b
2026-10-04 13:06 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-t15-gpu2lab_decode_c1_tps 9.50 tok/s · lab_prefill_tps 201 tok/sgatesT20b
2026-10-04 12:49 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012lab_decode_c1_tps 11.8 tok/s · lab_prefill_tps 218 tok/sgatesT21
2026-10-04 12:39 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012decode_window_min_tps 11.0 tok/s · tool_calls 24 callsT21
2026-10-04 12:16 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-g210-gpu012lab_decode_c1_tps 11.8 tok/s · lab_prefill_tps 219 tok/sgatesT21
2026-10-04 11:39 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-fitt2048-gpu012lab_decode_c1_tps 12.2 tok/s · lab_prefill_tps 153 tok/sgatesT21
2026-10-04 10:42 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012lab_decode_c1_tps 10.4 tok/s · lab_prefill_tps 132 tok/sgatesT21
2026-10-04 10:31 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012decode_window_min_tps 9.86 tok/s · tool_calls 22 callsT21
2026-10-04 09:30 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq3-xxs-128k-3x-ub2048-gpu012lab_decode_c1_tps 10.5 tok/s · lab_prefill_tps 126 tok/sgatesT21
2026-10-04 09:17 AEST✕ crashGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012lab_decode_c1_tps unavailable · lab_prefill_tps unavailableharnessT21
2026-10-04 09:07 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012decode_window_min_tps 11.4 tok/s · tool_calls 26 callsT21
2026-10-04 08:27 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-3x-ub2048-gpu012lab_decode_c1_tps 12.3 tok/s · lab_prefill_tps 158 tok/sgatesT21
2026-10-04 08:17 AEST! failFlash-Nextstabilityllama.cpp-qwen3.8-flash-next-iq4xs-128k-3x-fit-gpu012decode_window_min_tps unavailable · tool_calls 0 callsT11
2026-10-04 06:02 AEST✕ crashGLMstabilityik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2decode_window_min_tps unavailable · tool_calls 0 callsmeasureT22b
2026-10-04 05:33 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-ckpt4-gpu2lab_decode_c1_tps 8.60 tok/s · lab_prefill_tps 147 tok/sgatesT22b
2026-10-04 05:16 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2lab_decode_c1_tps 9.50 tok/s · lab_prefill_tps 199 tok/sgatesT20b
2026-10-04 05:06 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2decode_window_min_tps 8.85 tok/s · tool_calls 30 callsT20b
2026-10-04 04:26 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq2-m-128k-1x-ub2048-gpu2lab_decode_c1_tps 9.50 tok/s · lab_prefill_tps 200 tok/sgatesT20b
2026-10-04 03:40 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2lab_decode_c1_tps 8.40 tok/s · lab_prefill_tps 71 tok/sgatesT20b
2026-10-04 03:30 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2decode_window_min_tps 7.65 tok/s · tool_calls 25 callsT20b
2026-10-04 02:26 AEST! failGLMlabllama.cpp-glm-5.3-flash-iq3-xxs-128k-1x-gpu2lab_decode_c1_tps 8.30 tok/s · lab_prefill_tps 65 tok/sgatesT20b
2026-10-04 01:51 AEST✕ crashGLMstabilityik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2decode_window_min_tps unavailable · tool_calls 0 callsmeasureT22b
2026-10-04 01:19 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-1x-mtp-gpu2lab_decode_c1_tps 8.20 tok/s · lab_prefill_tps 147 tok/sgatesT22b
2026-10-04 01:01 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2lab_decode_c1_tps 9.50 tok/s · lab_prefill_tps 151 tok/sgatesT22b
2026-10-04 00:51 AEST! failGLMstabilityik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2decode_window_min_tps 9.46 tok/s · tool_calls 0 callsT22b
2026-10-04 00:22 AEST! failGLMlabik_llama-glm-5.3-flash-iq2-m-128k-1x-gpu2lab_decode_c1_tps 10.0 tok/s · lab_prefill_tps 151 tok/sgatesT22b
2026-10-03 23:31 AEST! failGLMlabllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2lab_decode_c1_tps 9.30 tok/s · lab_prefill_tps 80 tok/sgatesT20b
2026-10-03 23:21 AEST! failGLMstabilityllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2decode_window_min_tps 8.80 tok/s · tool_calls 19 callsT20b
2026-10-03 21:59 AEST! failGLMlabllama.cpp-glm-5.3-flash-q2-k-128k-1x-gpu2lab_decode_c1_tps 9.20 tok/s · lab_prefill_tps 78 tok/sgatesT20b
2026-10-03 12:48 AEST! failGLMlabglm-llamacpp-iq2m-128k-1x-gpu2lab_decode_c1_tps 9.60 tok/s · lab_prefill_tps 81 tok/sgatesT20b
2026-10-03 12:38 AEST! failGLMstabilityglm-llamacpp-iq2m-128k-1x-gpu2decode_window_min_tps 8.91 tok/s · tool_calls 24 callsT20b
2026-10-03 11:41 AEST! failGLMlabglm-llamacpp-iq2m-128k-1x-gpu2lab_decode_c1_tps 9.60 tok/s · lab_prefill_tps 81 tok/sgatesT20b
2026-10-03 11:03 AEST! failFlash-Nextstabilityfn-llamacpp-iq4xs-128k-1x-fit-gpu2decode_window_min_tps 23.8 tok/s · tool_calls 49 callsT11
2026-10-03 10:28 AEST✕ crashFlash-Nextloadfn-llamacpp-iq4xs-128k-1x-gpu2load_seconds unavailableloadT11
2026-10-03 08:06 AEST! failFlash-Nextstabilityfn-trellis-3.05-0xsero-main-x8-gpu2decode_window_min_tps 38.1 tok/s · tool_calls 63 callsT10

Decisions

0001-flashnext-llamacpp-quant.md 0001 — Flash-Next llama.cpp comparison quant (T11)

0001 — Flash-Next llama.cpp comparison quant (T11)

Decision: use bartowski/Qwen3.8-Flash-Next-GGUF@928589fdb66c6ff07f22ac561e3fbce76553548f IQ4_XS (97.7 GB) for both the 1-card and the 3-card llama.cpp comparison, on the attested ghcr.io/ggml-org/llama.cpp:server-cuda12-b11312 image. N-gram table (per_layer_token_embd.weight) pinned to CPU with -ot; experts on CPU for the layers that do not fit (--n-cpu-moe N, tuned once per layout), fp16→q8_0 KV only if needed for 128k+.

Why: closest to 4-bit that fits 24 GB VRAM + ~110 GiB RAM on one card (IQ4_XS ≈ 91 GiB); same file serves the 3-card run, so the 1 vs 3 comparison changes only the layout. Q4_K_M (119.6 GB ≈ 111 GiB) would not leave room on 1 card. No MTP (llama.cpp Flash-Next MTP PR #27836 not merged). Not REAP-pruned.

Limits: the kit quality panel speaks the SGLang API only, so this route gets no kit band → it cannot be the Flash-Next headline (SPEC §5); it is a comparison point. Delete after T11 unless it beats trellis.

Outcome after review (GPT-6.1 Sol: AMEND) — accepted
  • IQ4_XS stays the paired 1 vs 3 card file. Q4_K_M only as a single feasibility screen if IQ4_XS is competitive with trellis (disk is the binding constraint).
  • The 1 vs 3 pair freezes occupied context, KV format, batching and workload; --n-cpu-moe / tensor split per layout are reported, and any KV change is a separate comparison.
  • Full digest + shard hashes via tools/models.py; tensor placement of per_layer_token_embd.weight verified from the server log.
  • Eligibility experiment-only (no kit band). Results kept regardless; weights deleted only after manifests exist.
0001-flashnext-llamacpp-quant.review.md 0001-flashnext-llamacpp-quant.review
  • AMEND. Keep IQ4_XS as the primary paired comparison, but replace the categorical Q4_K_M rejection with a memory-budget check. The strongest counter-argument is that the decision compares the entire file against host RAM, despite offloading weights to VRAM.
  • Q4_K_M is not demonstrably excluded. 119.6 GB ≈ 111.4 GiB. If 18–20 GiB of weights reside on the GPU and their host pages are reclaimable, host weight residency becomes roughly 91–93 GiB, leaving 17–19 GiB within the stated 110 GiB RAM budget. Loading peaks, pinned allocations, KV and scratch could consume that margin; the supplied numbers do not prove they will.
  • Concrete alternative: use IQ4_XS for T11, with one bounded Q4_K_M feasibility screen at the intended context if the allocation budget supports it. Keep any Q4 results as a separate quant comparison. Do not expand into another tuning grid.
  • Fitting is not evidence of adequate speed. Dual-channel DDR4-3200 has a theoretical 51.2 GB/s ceiling. At 50 tok/s, that permits only 1.024 GB of DRAM traffic per token. If all 6B active LM parameters required fresh ideal 4-bit reads, that is ~3 GB/token and at most 17 tok/s before overhead. Record CPU offload placement and measured decode; this is a conditional bandwidth screen, not a performance prediction.
  • “Only the layout changes” needs qualification. Independently tuning --n-cpu-moe N compares each layout’s selected configuration, rather than isolating card-count scaling. Freeze occupied context, KV format, batching and workload across the pair; report both CPU-expert allocations and tensor splits. Treat any KV-format switch as a separate comparison.
  • Complete provenance before launch. Pin the full image digest already supplied in SPEC §6, not just server-cuda12-b11312. Add the exact GGUF shard list and hashes, immutable launches, and the actual tensor-placement evidence for per_layer_token_embd.weight. The proposed -ot setting remains unverified in this review.
  • The quality limitation is correctly identified, but applies to registry eligibility too. Without the required kit band, this Flash-Next configuration is experiment-only, even if all registry lab gates pass. Neither IQ4_XS nor a speed win establishes quality equivalence.
  • Replace “delete unless it beats trellis.” Define the comparison at matched occupied context and cache state, and retain results regardless of outcome. Delete rejected weights only after hashes and immutable references are recorded; retain any finalist under §9. A useful three-card result need not beat a single-GPU engine to justify publication.
0002-glm-quant-plan.md 0002 — GLM-5.3-Flash quant plan for llama.cpp / ik_llama.cpp (T20–T22)

0002 — GLM-5.3-Flash quant plan for llama.cpp / ik_llama.cpp (T20–T22)

Source: bartowski/GLM-5.3-Flash-BF16-GGUF@66e5e9f0e337470e8801b99bae6c45779b042cc2 (made with llama.cpp b11279, mainline-compatible; MTP layers kept at Q4_0). Unsloth GGUFs excluded: not loadable on mainline until re-uploaded.

LayoutQuantsSizeFit reasoning
1 card (24 GB VRAM + ~115 GiB usable RAM)IQ2_M, Q2_K120.6 / 125.7 GB (112 / 117 GiB)~21 GiB on GPU, rest in RAM
1 card, stretchIQ3_XXS138.8 GB (129 GiB)over RAM by ~10 GiB → mmap paging from NVMe; run once, expect slow, record it
3 cards (72 GB VRAM + RAM)IQ3_XXS, IQ4_XS138.8 / 177.7 GBIQ4_XS ≈ 165 GiB: ~66 GiB VRAM + ~99 GiB RAM; ~4-bit like every accepted GLM recipe

Order (disk is ~750 GB for everything, so sequential): IQ2_M → Q2_K → (IQ3_XXS) → delete the losing 2-bit → IQ4_XS. ik_llama.cpp + MTP runs on the same files.

Quality judged per SPEC §5 (tools gate, 0 parse failures over ≥100 calls, ≥80 % of reference solve rate in the uncapped eval).

Outcome after review (GPT-6.1 Sol: AMEND) — accepted
  • Order: IQ2_M and Q2_K on 1 card → same files on 3 cards → IQ3_XXS on 1 card (tight resident candidate, not presumed paging) and 3 cards → IQ4_XS on 3 cards.
  • Fit is checked by measured placement at the stated context, per GPU; uneven splits watched.
  • Nothing is a finalist before the uncapped quality evaluation; deletion only of rejected candidates, after manifests.
  • MTP on/off is its own variable family; MTP use verified from engine logs and metrics.
  • Publish tool-call counts and reference solve counts next to every quality verdict.
0002-glm-quant-plan.review.md 0002-glm-quant-plan.review
  • AMEND. Keep the candidate quants, but correct the IQ3_XXS fit claim and make deletion depend on qualification. The strongest counter-argument is that weight capacity does not establish useful speed or quality: the spec already estimates one-card GLM at 8–13 tok/s, below the 15 tok/s registry floor.
  • IQ3_XXS is not automatically a paging experiment. Using the decision’s assumptions: 138.8 GB ≈ 129.3 GiB; subtracting 21 GiB GPU residency leaves 108.3 GiB in RAM, about 6.7 GiB below the stated 115 GiB usable budget. Treat it as a tight resident-memory candidate. Declare paging only after accounting for context, runtime allocations and actual tensor placement.
  • All fit claims need a specified context and measured placement. Approximate host-weight budgets are 91.3 GiB for IQ2_M, 96.1 GiB for Q2_K, and 99.5 GiB for three-card IQ4_XS. Their remaining host margins are roughly 23.7, 18.9 and 15.5 GiB respectively. These exclude additional runtime memory. Likewise, allocating 66 GiB across three cards leaves only 2 GiB per card on average; an uneven split can exhaust one GPU.
  • Budget bandwidth before promising 15 tok/s. Dual-channel DDR4-3200 has a theoretical maximum of 51.2 GB/s. At 15 tok/s, that permits at most 3.41 GB of DRAM traffic per generated token, with a smaller practical allowance. Measure effective bandwidth, expert placement and decode; VRAM capacity alone cannot show that the one-card route clears this constraint. Also, Q2_K is only 4.2% larger than IQ2_M—kernel efficiency and quality could easily determine the winner.
  • “~4-bit like accepted recipes” is precedent, not qualification. The spec’s reported sub-4-bit KLD of 0.28–0.38 makes quality a substantial risk. Establish the reference explicitly: exact EXL3 if loadable, otherwise the highest-bit GGUF that fits. If that is IQ4_XS, retain it through all candidate evaluations. A speed winner cannot be called a finalist before the uncapped quality evaluation.
  • Concrete alternative order: screen IQ2_M and Q2_K on one card, compare those files on three cards, then screen IQ3_XXS on both layouts and IQ4_XS on three cards. Run uncapped quality evaluation for promising configurations, then qualification for survivors. Give one-card IQ3_XXS an ordinary residency check before labelling it “stretch”; make three-card IQ3_XXS an explicit comparison point.
  • Replace “delete the losing 2-bit” with “delete a rejected candidate.” All four listed quants total 562.8 GB, leaving about 185 GB of the specified 748 GB free before Qwen weights, images and other artifacts. Build the combined inventory first. Preserve hashes and file references before deletion, and retain qualifying finalists under §9; being slower than another quant does not itself establish rejection.
  • Shared files do not establish shared MTP support. Verify that each pinned engine loads this exact revision and actually uses the retained MTP layer. Compare MTP off/on as a separate variable family; record accepted tokens/s, acceptance rate, settings and quality. Keep locally built ik_llama.cpp results experiment-only until its image provenance becomes eligible.
  • The decision’s quality summary omits essential acceptance conditions. Require the rolling 60-second decode floor, three qualification runs, cold-cache reset, no OOM/crash/repetition, and all six lab gates. Zero failures in 100 tool calls satisfies the spec but still gives an approximately 3% one-sided 95% upper bound on failure probability. The solve-rate test is also coarse: with four reference solves, all four are required; with five, four suffice. Publish these counts alongside occupied context and exact launch provenance.
0003-harness-interpretations.md 0003 — Harness interpretations of SPEC §5 (repetition, parse failures)

0003 — Harness interpretations of SPEC §5 (repetition, parse failures)

Context: the T06a harness was exercised for 10 min against the operator's daily Qwen3.8-27B (EXL3 3 bpw, SGLang): 73 tool calls, 0 malformed tool calls, but 12 repetition hits — genuine reasoning loops (the same multi-paragraph block 3× inside one reasoning field) on sympy-21612 at 60–83k occupied tokens. Under §5 as written, the operator's everyday model fails stability.

Interpretations implemented by the harness:

  1. Repetition, per field: the 64-token ×3 rule is evaluated within each generated field (reasoning, content, each tool call's arguments), not over the concatenation, because code drafted in reasoning and then emitted in a tool call trips the joined check without being a loop.
  2. Parse failures, strict: besides invalid tool-call JSON and unparsed tool-call markup, every agent format error counts — including a response truncated at max_tokens (8192) with no tool call. Each response counts once; causes are split in detail.
  3. Same-call streak (≥4 identical consecutive tool calls) resets at each task attempt.

Proposed (pending data, operator decides): keep §5 strict for this first pass and record all hits with causes; after T10/T20 produce real data, revisit whether "no repetition loops" should mean "no unrecovered loop" (a loop that ends the response by hitting max_tokens or ends the task) — not changed now.

Outcome after review (GPT-6.1 Sol: AMEND) — accepted
  • SPEC §5 amended: per-field repetition (cross-field diagnostic only), detector frozen for T10/T20, streak resets only on a fresh conversation (each task attempt starts a new mini-swe-agent conversation, so the harness already complies).
  • Three error categories: parse_failures (malformed tool calls only — the qualification criterion), output_truncations, agent_format_errors, all published.
  • No relaxation of "no repetition loops" now. Recovery and wasted-token figures to be published as supplementary data; any protocol revision later is prospective and does not re-grade earlier results.
  • Denominators always shown next to zero-failure claims.
0003-harness-interpretations.review.md 0003-harness-interpretations.review
  • AMEND. Keep the strict no-loop gate, but distinguish detector clarification from changes to eligibility. The strongest objection is that the harness simultaneously narrows repetition detection and broadens parse failures, while claiming to preserve §5.
  • Per-field repetition is defensible, but requires an explicit spec amendment. “Within one response” currently includes repetitions across fields. Adopt: “Within reasoning, content, or each tool-call argument field, independently, using the served model’s tokenizer.” Retain cross-field matches as diagnostics so the excluded cases remain auditable.
  • Do not label every agent-format error a tool-call parse failure. Invalid JSON and malformed tool markup qualify; an otherwise valid response exhausted at max_tokens without a tool call is a truncation or agent-protocol failure. Record tool_parse_failure, agent_format_error, and output_truncation separately. If all must disqualify, add that requirement explicitly to §5.
  • Resetting the same-call streak is appropriate only at an independent task attempt. Specify that a fresh agent conversation resets it; a retry continuing the same conversation does not. Otherwise attempt boundaries can conceal four consecutive identical calls.
  • The daily-model result does not justify relaxing the gate. Twelve genuine loops establish that this particular screening run fails the agreed criterion. They do not establish that the criterion is unsuitable for the two target models. The observed 60–83k occupancy also falls below the ≥128k requirement for “agentic-usable.”
  • Reject “unrecovered loops only” as the qualification alternative. Recovery does not erase wasted generation or latency. At exactly the specified decode floors, an 8,192-token response takes approximately 164 seconds for Flash-Next or 546 seconds for GLM, excluding prefill—27% or 91% of a ten-minute run. Publish recovery and wasted-token measurements as supplementary metrics while retaining zero detected loops for qualification.
  • The parse evidence is preliminary. Seventy-three calls falls short of the ≥100-call minimum. Under an independent binomial assumption, 0/73 failures still gives a one-sided 95% upper failure-rate bound of approximately 4.0%; 0/100 gives 3.0%. Report the denominator without presenting zero observed failures as strong reliability proof.
  • Concrete alternative: ratify the field boundary and separate error categories before target comparisons; freeze that detector version for T10/T20; preserve each offending response and count affected responses separately from detector matches. Use later data to propose a prospective protocol revision, with existing results retaining their original qualification status.
0004-core2-quarantine.md 0004 — Quarantine CPU core 2 (logical CPUs 2 and 18)

0004 — Quarantine CPU core 2 (logical CPUs 2 and 18)

Evidence: two hard resets on 2026-10-03 (04:57 near-idle, 06:56 during T10 model load). Each next boot logged an uncorrectable machine check in bank MC5 (Execution Unit) — CPU:18 MC5_STATUS[-|UE|MiscV|-|PCC|TCC|SyndV] and CPU:2 MC5_STATUS[-|UE|MiscV|AddrV|PCC|TCC|SyndV]. CPUs 2 and 18 are SMT siblings of physical core 2. No OOM, no GPU Xid, no kernel panic text: the core faulted and the platform reset. Likely cause: per-core Curve Optimizer/PBO undervolt too aggressive for that core, or core degradation (RMA case if BIOS is stock; BIOS 5101).

Decision (operator, option 2): take CPUs 2 and 18 offline for the whole campaign (resets at reboot; hygiene.sh check refuses to benchmark while they are online). Every record carries host.cpus_online/cpus_offline. llama.cpp and ik_llama launches use 15 threads. Then validate with tools/stability/cpu_check.sh (all-core AVX2, per-core single-thread boost bursts, all-core openssl) and look for new machine checks.

Consequence for results: numbers are for a 15-core 5950X. CPU-bound routes (GLM experts on CPU) are affected ~6 %; GPU/PCIe-bound trellis decode should not be. Any result produced before this decision does not exist (T10 never completed).

Check result (2026-10-03 07:25–07:40)

cpu_check.sh 300 20 180 on the 30 online CPUs: 0 machine checks, no reset (log: logs/cpu_check-core2-offline.log). Not conclusive for the near-idle failure mode; the hygiene check keeps the quarantine enforced for every run.

Revision (2026-10-03 10:30) — quarantine lifted with a tripwire
  • Operator BIOS check: PBO was Auto (ASUS stock); Curve Optimizer set to all cores, positive 0 (= no offset). So the CPU was effectively at stock when it faulted; a weak core 2 remains the leading hypothesis.
  • TARGET_CPUS="2 18" cpu_check.sh 300 20 180 with all 32 CPUs online: all-core AVX2, per-core bursts, 5 min sustained on CPU 2 and 18 each, 5 min on/off bursts on each, all-core openssl — 0 machine checks, no reset.
  • Not conclusive (both crashes came after hours of uptime). Core 2 is back online for full capacity; hygiene.sh check refuses to benchmark and re-offlines CPUs 2,18 if any machine check appears in the current boot; every record carries host.machine_checks_this_boot. llama.cpp/ik_llama launches back to 16 threads.
  • T10 ran on 15 cores (recorded in its host block). Trellis decode is PCIe-bound; one 16-core re-check is planned.
  • If it recurs at stock: keep core 2 offline for the campaign; RMA candidate.
Revision (2026-10-04 08:05) — permanent quarantine
  • Third fatal reset with core 2 online at stock: 2026-10-04 06:14, ~20.5 h into the boot, during the 16-core T10 re-check model load (CPU:2 MC5_STATUS[-|UE|MiscV|AddrV|PCC|TCC|SyndV], logged at the 06:54 boot). Same pattern as the 2026-10-03 06:56 crash (Flash-Next load) after a long uptime.
  • Core 2 stays offline for the rest of the campaign, enforced at boot by local-ai-bench-quarantine.service (tools/box/install.sh) and by hygiene.sh check (CPU_QUARANTINE defaults to "2 18"). All llama.cpp/ik_llama launches use 15 threads. The 16-core T10 re-check is dropped; all results are for a 15-core 5950X.
  • Stock settings + repeated MC5 UE on one core: AMD RMA candidate (operator's call).
0005-ik-llama-glm-tool-parsing.md 0005 — ik_llama.cpp does not parse GLM-5.3-Flash tool calls (commit 5bf8f0fe)

0005 — ik_llama.cpp does not parse GLM-5.3-Flash tool calls (commit 5bf8f0fe)

Finding (T22b, IQ2_M, 1 card): the model emits well-formed GLM tool calls (<tool_call>bash<arg_key>command</arg_key><arg_value>…</arg_value></tool_call>) but ik_llama's server returns them as plain content with empty tool_calls: lab tools gate fails; stability run 0 tool calls / 47 parse failures (unparsed_markup). ik_llama builds tool parsers from the chat template (common/chat-auto-parser*); mainline llama.cpp b11312 parses the same GGUF correctly (tools gate passed in T20b). No matching ik_llama issue/PR found (search 2026-10-04).

Speed is unaffected and recorded: decode 10.0 tok/s (mainline 9.6), prefill 151 lab / 203 TTFT-8k (mainline ~80).

Decision: keep the ik_llama route as speed evidence only (experiment-only, fails §5 tools); do not patch ik_llama in this campaign — single-card GLM is ~10 tok/s on every route, below the 15 tok/s gate, so tool parsing would not change the outcome. Phase 3 option: upstream bug report with the captured response (artifacts/runs/…ik_llama…stability/proxy/parse_failures/).

3-card follow-up (2026-10-04, GPU2-first order, decision 0006)
  • No MTP, --fit-margin 4096: lab decode 13.4 / 14.0, agentic 13.3 tok/s steady (best GLM decode on this box), prefill 185–283; tool calls still unparsed (68 agent-format errors), tools gate fails.
  • MTP, --fit-margin 7168 (4096 OOMed creating the MTP context): lab decode 10.2 (acceptance 51 %), agentic 15.6 whole-run before the server crashed mid-run. The extra margin for the draft context pushes experts to host RAM.
  • Route parked: experiment-only speed evidence. The most interesting number for 0xSero is 13–14 tok/s without MTP; a tool-parser fix in ik_llama plus an attested image would be needed before it could be a 3-card candidate.
0006-pcie-layout.md 0006 — PCIe lane layout: what can be reconfigured, and what it buys

0006 — PCIe lane layout: what can be reconfigured, and what it buys

Operator question (2026-10-04): can the PCIe lanes be reconfigured to get better model performance?

Measured topology (lspci, nvidia-smi, T05a)
DeviceAttachLinkMeasured H2D
GPU0 3090 04:00.0X570 chipset, behind a Gen4 x4 uplink shared with the KC3000 1 TB NVMe (root + all model files), NIC, BMC, SATA, USBGen4 x8 to the chipset6.0 GB/s
GPU1 3090 0b:00.0CPU lanes, slot 2 (x8/x8 split) — drives the displayGen4 x8not measured (display)
GPU2 3090 0c:00.0CPU lanes, slot 1 (x8/x8 split)Gen4 x813.5 GB/s
990 EVO Plus 4 TBCPU lanes (M.2_1, Gen4 x4)unused: holds another OS install

Already optimal and not a lever: every link trains at Gen4 (16 GT/s); Resizable BAR is on (32 GB BAR1); ASPM is disabled on every GPU link. AM4 has 24 usable CPU lanes (16 slots + 4 M.2 + 4 chipset uplink); bifurcation or risers cannot add lanes, so the only real choices are which card sits where and how many cards share the 16 slot lanes.

Where PCIe bandwidth limits each route (evidence from our runs)
  1. Flash-Next trellis (experts streamed from host RAM to one GPU): decode is PCIe-bound. 0xSero's data: ~26 GB/s → 62.7 tok/s, 14–16 GB/s → 53.2. Ours at 13.5 GB/s: lab 49–55, kit C1 46.9, agentic 43. x16 should add ~15–25 %.
  2. llama.cpp/ik_llama prefill with experts in RAM is PCIe-bound (CPU-resident weights are copied to a GPU for large batches). Evidence: on GPU2, raising -b/-ub 512 → 2048 lifted GLM IQ2_M prefill 85 → 200 tok/s (each copy serves 4× more tokens). And 3 cards prefilled slower than 1 card (158 vs 200), consistent with the copy going to CUDA0 = GPU0, the 6 GB/s chipset card.
  3. llama.cpp decode with experts on the CPU is DDR4-bound, not PCIe-bound (GLM ~9.5 tok/s on 1 card, 12.3 on 3 cards, from less RAM-resident weight). PCIe changes nothing here.
Options
#ChangeCostExpected effect
S13-card llama.cpp with the CPU-attached GPU2 as CUDA0 (CUDA_VISIBLE_DEVICES=2,1,0 in the container)none3-card prefill back to ≥ 1-card level (~200+), no decode change
S2Larger ubatch (-b/-ub 4096) on llama.cpp routesnone (more VRAM for compute buffers)prefill up to ~1.5–2× again; watch VRAM (the IQ2_M 3-card run already OOMed on CUDA0)
P1One card alone in slot 1 at x16 (T12, already planned)operator moves cardsFlash-Next decode +15–25 % (the only route to ≥ 50 headline at full context); llama.cpp prefill ~2× vs x8; GLM exact mode at the bandwidth its 12.6 tok/s reference used
P2Plug the monitor into GPU0 (chipset card) instead of GPU1move one cableGPU1 + GPU2 (both CPU x8) free for 2-card runs without going headless; desktop stays usable during single/2-card work
P3Model files on the CPU-attached 990 EVO Plus instead of the chipset KC3000needs the operator's OK (other OS lives there)removes model loads and Flash-Next n-gram NVMe reads from the chipset uplink that GPU0 also uses; small effect except for 3-card runs with GPU0
P42 cards on CPU lanes (x8/x8), third card removed or idlenone (GPU1+GPU2)avoids the 6 GB/s card; less VRAM (48 GB) so more experts in RAM
Decision (proposed)
  • Run S1 and S2 now in the current layout (software only, attributable, one variable family each), on GLM IQ2_M: 3-card GPU2-first at ub2048 (vs the existing GPU0-first record), then 1-card GPU2 at ub4096 (vs ub2048).
  • Keep P1 (T12) as the main hardware lever; in the x16 session also rerun the best GLM llama.cpp 1-card config so the x8→x16 effect on prefill is measured directly.
  • Recommend P2 to the operator at the T12 visit (same visit, one cable). P3 only if the operator wants to free the 990 EVO Plus; not worth touching another OS for this campaign.
  • Publish this as a site methodology section with the before/after records.
Outcome after review (GPT-6.1 Sol: AMEND) — accepted, plus operator rule
  • Bandwidth effects are hypotheses until measured. Expected-gain numbers above are estimates, not claims; x16 is the principal hardware hypothesis for Flash-Next, not a guarantee of qualification.
  • S1 is a fresh paired comparison, both headless, identical except card order: IQ2_M 3-card ub2048 with a 2 GiB --fit-target margin (the default-margin run OOMed on CUDA0 after 62 min), GPU0-first vs GPU2-first (CUDA_VISIBLE_DEVICES=2,1,0; the container sees the cards in PCI order). Every TTFT step now logs per-GPU PCIe rx/tx (nvidia-smi dmon -s t) to show which card receives the weight copies. Decode changes are reported, not assumed.
  • S2: 1-card GPU2 ub2048 vs ub4096, both at 15 threads (the old ub2048 record ran 16 threads), peak VRAM and OOMs recorded.
  • P1: rerun the best 1-card GLM llama.cpp config at x16 with the same bandwidth procedure (x8 control exists).
  • Operator rule (2026-10-04): GPU0 is used last. Single-card runs use GPU2; 2-card runs GPU2+GPU1; 3-card runs order the cards GPU2, GPU1, GPU0. Pending S1 evidence, all new multi-card launches use the GPU2-first order.
  • P2 offered to the operator at the T12 visit; P3 deferred (would need a §9 exception and touches another OS).
S1 result (2026-10-04 12:15–13:05 AEST) — card order, GLM IQ2_M 3-card ub2048, 2 GiB fit margin, headless
GPU0 first (control)GPU2 first (CUDA_VISIBLE_DEVICES=2,1,0)change
TTFT prefill 8k / 32k159 / 169 tok/s240 / 242 tok/s+51 % / +43 %
lab prefill153 tok/s218 tok/s+43 %
lab decode C112.2 tok/s11.8 tok/s−3 %
PCIe rx during TTFT (avg / peak)GPU0 3.9 / 6.5 GB/s (link-saturated), others ~0.3GPU2 5.7 / 14.3 GB/s, others ~0.5

The host→device weight copies for prefill go to CUDA0; putting a CPU-lane card there lifts prefill ~1.5×. Decode is within noise of unchanged (slightly lower). The GPU2-first config also finished a clean 10-min agentic run (11.3 tok/s, 24 tool calls, 0 parse failures, 0 loops) and passed the post-run gates without the OOM that hit the default margin. Rule confirmed: GPU0 last.

S2 result (2026-10-04 13:05–14:23 AEST) — ubatch, GLM IQ2_M 1 card GPU2, 15 threads
ub2048 (control)ub4096change
lab prefill201 tok/s288 tok/s+43 %
TTFT prefill 8k / 32k213 / 222321 / 316+51 % / +42 %
lab decode C19.59.6unchanged
VRAM at ready21.9 GB22.9 GB+1 GB, no OOM in a 10-min agentic run

The 15-thread control matches the earlier 16-thread ub2048 record (9.5 / 200), so the core-2 quarantine costs nothing measurable on this route. Gain > 3 %, so the search continues: ub8192 on 1 card, and ub4096 on the GPU2-first 3-card config.

S2 step 2 (2026-10-04 16:13–16:30 AEST) — ub8192, 1 card GPU2

lab prefill 370 (+28 % vs ub4096), TTFT 384 / 404, but lab decode 9.1 (−5 %): the larger compute buffer makes --fit keep more experts in host RAM (VRAM at ready 22.2 GB). Prefill and decode now trade off, so the batching search stops here: ub4096 is the balanced setting (decode unchanged, prefill +43–51 %); ub8192 is the prefill-maximising variant, published as such.

  • 3-card GPU2-first, ub2048 → ub4096 (16:30–16:58): lab prefill 218 → 258 (+18 %), TTFT 240 → 298 (+24 %), lab decode 11.8 → 11.2 (−5 %). Same prefill/decode trade-off as 1 card at ub8192; both published, no further steps.
0007-glm-current-layout-finalists.md 0007 — GLM current-layout finalists (T25a) and where the quality evaluation runs

0007 — GLM current-layout finalists (T25a) and where the quality evaluation runs

Facts (all current-layout GLM records, 2026-10-03/04)
RouteCardsLab decodeLab prefill10-min run (tool calls / parse failures / loops / solved)
llama.cpp IQ2_M ub409619.628827 / 0 / 1 / 1 of 2
llama.cpp IQ2_M ub2048, GPU2 first311.821824 / 0 / 0 / 1 of 2
llama.cpp IQ2_M ub2048, GPU0 first312.315826 / 0 / 0 / 1 of 2 (server later OOMed)
llama.cpp IQ3_XXS ub2048310.413222 / 0 / 1 / 1 of 2
llama.cpp IQ4_XS ub2048, GPU2 first38.315625 / 0 / 0 / 1 of 2
llama.cpp Q2_K / IQ3_XXS19.2 / 8.378 / 650–1 solved
ik_llama IQ2_M (± MTP)18.2–10.0 (MTP agentic ~12.6)~1500 tool calls: GLM tool calls not parsed (0005)

Every route is below the 15 tok/s registry gate on 1 card and on 3 cards (best 12.3). Decode is DDR4-bound: more cards only help by holding more experts in VRAM. Pending in the queue: IQ2_M 1-card ub8192 and 3-card ub4096 (prefill-only), ik_llama 3-card ± MTP (speed evidence only).

Proposal
  1. No current-layout GLM finalist. §5 qualification requires decode ≥ 15 tok/s in every window, so the 3 × 10-min qualification runs would fail by construction. T25a records "no finalist: below the floor" for every current-layout config instead of running them. All configs stay published as documented experiments (§6 Phase 2).
  2. One uncapped quality evaluation, of the best mainline weights, in the x16 session. Solve rate depends on weights, engine and sampling, not on the PCIe layout; the layout only changes how long the evaluation takes. So the T07 candidate evaluation of llama.cpp IQ2_M runs on the single x16 card next to the T23 reference (exact mode, or the highest-bit GGUF that fits if exact mode cannot load), with the identical task list, sampling and 60k-token budget. Each evaluation is estimated at 4–7 h (measured ~3–4k completion tokens per 10 min).
  3. No uncapped evaluation for IQ3_XXS / IQ4_XS / Q2_K / ik_llama. They are slower than IQ2_M and cannot become candidates; their 10-min screening solve counts are published labelled "screening only". IQ4_XS only fits on 3 cards, so measuring it would cost ~6 h of current-layout time for an informational number.
  4. Consequence: T12 (card move) can happen as soon as the current queue finishes (~3.5 h), not after ~1 day of current-layout quality runs. After the queue, IQ3_XXS and IQ4_XS are marked rejected and deleted (manifests kept) to make room for the x16 weights (Flash-Next 2.05/4.05 bpw: 63/108 GB; GLM exact EXL3: 125 GB).
Outcome after review (GPT-6.1 Sol: AMEND) — accepted
  • Below-floor is measured, not inferred from lab C1: every current-layout llama.cpp config with a 10-min run has a measured rolling 60-s generation-only minimum below 15 (IQ2_M 3-card 11.0–11.4, IQ2_M 1-card 8.8–9.1, IQ3_XXS 3-card 9.9, IQ4_XS 3-card 7.8). Configs without a stability run (prefill-only S1/S2 controls, ub8192) are recorded as "qualification skipped: subfloor lab C1; rolling floor not measured". T25a closes with no current-layout finalist once the pending queue (ik_llama 3-card ± MTP, ub follow-ups) is in; ik_llama stays experiment-only (0005).
  • Quality evaluation is conditional. At x16: measure IQ2_M 1-card first; only a config that qualifies (§5 incl. the rolling floor, ≥ 100 cumulative tool calls, post-run gates) gets the uncapped T07 candidate evaluation. An optional quality study of a non-qualifying config is labelled informational. Quality results apply to the tested configuration; layout independence is not assumed.
  • Runtime estimate corrected: the 60k budget is per task; worst case 2.5–3.3 h per task, 12–33 h per evaluation (less when tasks submit early). Budget accordingly; this is another reason to evaluate qualifiers only.
  • Labels: IQ4_XS, IQ3_XXS (3-card), Q2_K: "rejected for speed, quality unmeasured". IQ2_M 3-card GPU0-first: OOM after 62 min (default fit margin). IQ2_M/IQ3_XXS repetition hits are published with the run.
  • Storage: keep IQ2_M (x16 rerun) and IQ3_XXS (highest-bit GGUF that fits 1 card = fallback reference if exact mode cannot load); delete IQ4_XS after the queue, manifests and hashes kept for reacquisition.
0008-flashnext-3card-llamacpp.md 0008 — Flash-Next on 3 cards with llama.cpp: tune it as a separate 3-card candidate

0008 — Flash-Next on 3 cards with llama.cpp: tune it as a separate 3-card candidate

Evidence (T11, 2026-10-04 15:38 AEST, IQ4_XS, GPU2,GPU1,GPU0 order, 2 GiB fit margin, -b 2048 / -ub 512 defaults)
  • Lab: all six gates pass, decode C1 57.1 / 56.9 tok/s (before / after the agentic run), lab prefill 608 / 605.
  • TTFT prefill: 987 tok/s at 8k, 852 at 32k.
  • 10-min agentic run: 70 tool calls, 0 parse failures, 0 loops, 1 of 2 solved, whole-run decode 45.4, lowest 60-s window 38.6 (fails the 50 floor; decode falls as the context grows).
  • VRAM at ready 65.5 GB of 72; weights 97.7 GB, so ~30+ GB (incl. the n-gram embedding table) stays in host RAM.
  • For comparison, the 1-card trellis/SGLang baseline at x8: lab 49–55, kit C1 46.9, agentic 43.

Per SPEC §6 a 3-card config is a separate, clearly labelled recipe; llama.cpp routes are judged by the registry lab and §5 (the kit's SGLang-only panel cannot score a GGUF, so it is "kit band not measurable").

Proposal (current layout, before T12; one variable family per step, stop rule 3 × < 3 %)
  1. Re-download IQ4_XS (deleted at 16:13 per plan; manifest kept, 97.7 GB, ~15 min).
  2. Prefill: -b/-ub 2048, then 4096 (GLM showed +135 % then +43 %; target lab prefill > 1000).
  3. Decode floor: the floor fails late in the run as context grows. One quant step down to IQ3_M (93.2 GB) or IQ3_XXS (88.0 GB) puts more experts in VRAM; run the better prefill setting from step 2 on it. Quality cost is unmeasured by the kit for GGUF; report solve counts and parse failures, label "quality: screening only".
  4. If a config reaches lab decode ≥ 50, lab prefill > 1000 and holds the 50 floor in a 10-min run, it becomes a 3-card finalist: §5 qualification (3 runs incl. cold and max context, ≥ 100 tool calls) before T12.
  5. Disk: the x16 downloads (296 GB) wait until this route is done and its rejected GGUFs are deleted.
Outcome after review (GPT-6.1 Sol: AMEND) — accepted
  • Route is experiment-only unless an accepted scoring path establishes the Flash-Next kit band for a GGUF; kit band recorded as unavailable, registry/headline eligibility unestablished. The T11 config is a failed screening config (floor −29.5 %, lab prefill −65 % vs targets), reported per constraint.
  • Batching steps, predeclared, judged by official lab prefill (decode and VRAM regressions checked): A -b 2048 -ub 2048, B -b 4096 -ub 4096, both vs the existing -b 2048 -ub 512 record. (GLM evidence says the ubatch sets the host→GPU copy amortisation, so b and ub move together.)
  • Before any quant change: decode vs occupied context from the existing T11 agentic trace (per-request generation speed against prompt tokens). IQ3_M is screened only if that shows a residency/transfer limit, not attention cost; its step is judged by the minimum rolling-window decode at matched occupancy. Expert/n-gram placement and memory use recorded from the server log.
  • Qualification only for a speed-passing config, naming its maximum occupied context (≥ 4k headroom), cold resets, post-run gates, ≥ 100 tool calls cumulative. Phase timebox kept; x16 is not deferred beyond this route's steps.
Decode vs occupied context (T11 3-card agentic trace, 55 requests, generation-only tok/s per request)
prompt tokensper-request decode
1.5k–6k (django task)35–51, typical ~42
6k–20k30–49, typical ~40
20k–31k30–41, typical ~38
31k–39k28–38, typical ~33

Two separate gaps to the 50 floor: a base gap (~42 tok/s already at 2–6k tokens, vs lab C1 57 measured over one long generation; agentic requests are short, 70–600 tokens) and a context slope of roughly −20 % from 5k to 39k. One quant step (IQ3_M, −4.6 % of bytes; IQ3_XXS −9.9 %) cannot close a 25–30 % gap even if every saved byte moved experts to VRAM. IQ3_M step dropped; the predeclared batching steps still run (prefill evidence, cheap). After them the route closes as experiment-only and its GGUF is deleted for the x16 weights.

Batching result (2026-10-04 18:00–18:23 AEST) — route closed
-b / -ublab decode C1lab prefillTTFT 8k / 32k
2048 / 512 (T11 record)57.1608987 / 852
2048 / 204844.3 (−22 %)717 (+18 %)1067 / 931
4096 / 409635.0 (−39 %)713980 / 881

Larger ubatches buy little prefill and cost a lot of decode: the bigger compute buffers make --fit move experts to host RAM. No setting reaches lab prefill > 1000 while keeping decode ≥ 50, and the agentic floor already failed at the default. Route closed: experiment-only, best config = the default-batch T11 record (all six gates, lab 57.1, agentic floor 38.6). GGUF deleted for the x16 weights (manifest kept).

0009-no-x16-session.md 0009 — No x16 session: Phase 1 and T23 continue in the current layout (operator decision)

0009 — No x16 session: Phase 1 and T23 continue in the current layout (operator decision)

Operator, 2026-10-04 19:15 AEST: skip the x16 test; the expected difference is judged marginal. T12 (card move) and T05b are cancelled. The operator's call stands; the estimate it overrides is recorded for the reader: 0xSero's reference gives +17.9 % decode for 14–16 → ~26 GB/s, and our x8 baseline (lab C1 49–55, kit C1 46.9, agentic 43, floor not held) sits just under the 50 tok/s target, so reaching it now depends on tuning at x8.

Re-plan (single card = GPU2, CPU-lane x8, 13.5 GB/s; GPU1 measured identical)
  1. T13 PR #144 rerun at x8 against image f77c72f6 (lab, kit sweep with early exit off, TTFT, panel, 10-min run). Proof for 0xSero is "x8 on this box"; whether he accepts it is asked in Phase 3.
  2. T15 brought forward: the main recipe on EXL3 2.05 and 4.05 bpw, only the weights changed. At x8 the expert bytes per token are the bottleneck, so 2.05 bpw is the most likely route to ≥ 50 sustained; the kit panel decides whether it stays in the quality band (required for the headline and registry PR). 4.05 bpw is the quality anchor.
  3. T23 GLM exact mode on GPU2 at x8 (route test + quality reference; may not load: ~119 GB pinned on ~125 GiB).
  4. T14a/T14b grid on the best quant from 1–2 (expert cache GB × context × KV format; then prefill chunk and n-gram tier), designed as its own decision once 1–3 are in. T16 qualification and headline follow on the winner.
  5. GLM IQ2_M x16 rerun dropped; IQ3_XXS kept as the fallback reference.
Outcome after review (GPT-6.1 Sol: AMEND) — accepted
  • Recorded as a deliberate restriction of the search, not evidence that x16 is marginal. Gap to target at x8: kit C1 46.9 → 50 needs +6.6 %; agentic 43 → 50 needs +16.3 %; there is currently no qualifying Flash-Next result.
  • Order: T13 at 3.05 → matched 2.05 / 4.05 screens (same launch, same context) → kit panel as a filter (top-1 ≥ 0.987, KL ≤ 0.0013) before any headline time → bounded tuning of in-band candidates, 3.05 included. Out-of-band results are published with kit eligibility, never as the registry path. "Bytes per token" is a hypothesis; transfer/cache evidence is recorded where the engine exposes it.
  • Acceptance unchanged: every eligible rolling 60-s window ≥ 50 and lab prefill > 1000 on the same config at the claimed occupied context. If no x8 config meets speed and quality together, that outcome is published as is.
  • T23 at x8: exact mode may be the GLM quality reference even below 15 tok/s (not a recipe); peak host memory and any load failure recorded (~110.8 GiB pinned vs ~125 GiB MemTotal). If it cannot load, the reference is the highest-bit GGUF that actually fits one card here, determined by a load test, not assumed to be IQ3_XXS.
  • Timebox and the 3 × < 3 % stop rule stay. Cancelled tickets (T12, T05b, x16 parts of T13/T14/T23) are recorded as skipped with the operator's reason.
0010-cpu-boost-off.md 0010 — CPU boost off for the rest of the campaign

0010 — CPU boost off for the rest of the campaign

Evidence (2026-10-04 19:48 AEST): hard reset during the T13 Flash-Next SGLang model load with core 2 already offline. The next boot decoded CPU:1 MC5_STATUS[-|UE|MiscV|AddrV|PCC|TCC|SyndV], "uncorrected error caused a data fabric sync flood" — the same bank, status and syndrome as the core-2 faults (0004), now on core 1. Tally of fatal resets: 5, of which 3 during a Flash-Next SGLang load (expert repacking: heavy multi-threaded host work) and 2 after long uptime. Temperatures at idle are normal (Tctl ~56–60 °C). A second core failing the same way points to the CPU (or its voltage under boost transients) rather than one bad core.

Decision: disable Precision Boost from Linux (/sys/devices/system/cpu/cpufreq/boost = 0, all cores capped at the 3.4 GHz base clock, lower voltage), applied at boot by local-ai-bench-quarantine.service and immediately on 2026-10-04 20:12. Core 2 stays offline. Every record now carries host.cpu_boost and host.cpu_max_mhz.

Consequence: CPU-bound work (GLM experts on CPU, model loads) gets slower; PCIe-bound trellis decode should not. Records before 20:12 ran with boost on and are labelled by that field's absence (= boost on). If faults continue with boost off, the CPU is not usable for unattended runs and the remaining work pauses for the operator (RMA).

0011-flashnext-x8-tuning-grid.md 0011 — Flash-Next tuning at x8 (T14): grow the expert cache

0011 — Flash-Next tuning at x8 (T14): grow the expert cache

Evidence
  • T15 is not runnable: the offload engine's HostExpertStore only supports 3-bit experts (offload_store.py: if bits != 3: raise NotImplementedError), so 2.05 and 4.05 bpw crash at load in both images. Only 3.05 bpw remains (the CPU-expert fallback, SGLANG_EXL3_MOE_OFFLOAD=cpu, is far slower and not pursued).
  • T13 (PR #144 launch, f77c72f6, x8, boost off): lab C1 54.5–55.6, lab prefill 2,433–2,596, kit C1 48.2 (32k: 49.9), kit panel in band (top-1 0.9887, KL 0.00097), agentic whole-run 44.1, min 60-s window 37.9 (floor 50 fails).
  • Agentic per-request decode is flat across context (median 43.7 / 43.8 / 41.7 tok/s at <8k / 8–20k / 20–40k), so context length is not the limiter; the gap to lab C1 (55) is consistent with expert-cache misses on coding traffic (each miss streams experts over x8 PCIe). Hypothesis, to be tested by changing the cache size.
  • VRAM budget at ready (server log): expert cache 11.08 GB (auto-fit: free 19.82 − staging 0.48 − reserve 8.26), KV pool 5-bit 1.77 GB for 210k tokens, 3.23 GB still free after CUDA graphs. Context length is therefore a weak lever (the whole KV pool is 1.77 GB); the reserve is the strong one.
Grid (base = PR #144 launch on the latest main-built image a21efea6, GPU2 x8; one variable per step)
StepChangeExpected expert cache
R1SGLANG_EXL3_EXPERT_CACHE_RESERVE_GB 8.26 → 6.0~13.3 GB (+20 %)
R2reserve → 5.0 (only if R1 is stable and gains ≥ 3 %)~14.3 GB (+29 %)
N1SGLANG_EXL3_NGRAM_TIER nvme → pinned (32.6 GB pinned RAM) on the best Rn-gram reads from RAM
G1--max-running-requests 1 --cuda-graph-max-bs-decode 1 (single-stream agentic) on the best, then re-fit the reservefrees graph memory

Metric per step: minimum rolling 60-s window and median per-request decode in the 10-min agentic run (the floor), with lab C1 and lab prefill checked for regressions. Every step runs lab,ttft,stab; the kit panel is re-run on the winner only (cache size does not change the math). Stop rule: 3 consecutive steps < 3 % on the floor metric. If no step lifts the minimum window to ≥ 50, Phase 1 closes with "not reached at x8", publishing the best config.

Outcome after review (GPT-6.1 Sol: AMEND) — accepted
  • Cache misses are a hypothesis: the floor needs +31.9 % (26.4 → 20.0 ms/token); a 20–29 % larger cache need not deliver that. Expert-cache hit/miss and transfer stats are recorded if the engine logs them, else marked unavailable.
  • Order: R0 (unchanged PR #144 launch on a21efea6) → R1 reserve 6.0 → G1 single-stream graphs on R1 → reserve refit only if peak VRAM (prefill and max-occupancy) shows room → N1 pinned n-grams with host-RAM headroom recorded. R2 (5.0) is not run blind: measured headroom after R1 would be ~0.97 GB, after R2 ~−0.03 GB.
  • max-running-requests 1 is a recipe concurrency constraint, stated as such.
  • Floor = minimum rolling 60-s generation-only window after the first 60 s (harness definition); window coverage reported. One 10-min run screens; the winner then needs full §5 qualification, kit band, and ≥ 128k occupied for "agentic-usable". Closure wording: "50 tok/s floor not reached in the tested x8 configurations".
R0 result and re-plan (2026-10-04 22:42 AEST)

R0 (a21efea6, PR #144 launch unchanged): load 260 s, lab C1 55.8, lab prefill 2,419, kit C1 47.4 (32k 48.6), TTFT 2,009 / 2,681, panel in band (top-1 0.9897, KL 0.00097) — then the 10-min agentic run crashed: CUDA OOM. After serving starts, Triton kernels load lazily (log: "device-loaded after serving started … free 0.83 → 0.09 GiB") and consume the 3.23 GB that looked free at ready. So the 8.26 GB reserve is needed on this image; shrinking it cannot grow the expert cache safely. (f77c72f6 survived the same agentic run at the same reserve.) Re-plan: R1 (6.0) is already running and is kept as evidence; R1+G1 dropped. Next: R0 repeat (is the OOM reproducible?), G1 at reserve 8.26 (single-stream frees graph and request memory → headroom), N1 pinned n-grams.

  • R1 (reserve 6.0, 22:42): load OK, then CUDA OOM at the first lab request — reserve cannot go below 8.26.
  • R0 repeat, agentic run only (22:58–23:22): no OOM; 71 tool calls, 0 parse failures, whole-run 43.3, min window 39.7; lab after 57.4 / 2,537. The first R0's OOM is intermittent (it followed lab, kit sweep, TTFT and panel on the same server). Reported as an a21efea6 stability risk for PR #144; frequency unknown from 2 runs.
  • G1 single-stream at reserve 8.26 (23:22): min window 40.1 (+1 % vs R0 repeat), whole-run 43.1, lab 54.6–55.6.
  • N1 pinned n-grams (23:48): min window 40.9 (+3.1 %), whole-run 44.6, lab C1 53.7 / 51.3 (lower), prefill unchanged.
Closure (2026-10-05 00:11 AEST)

Stop rule met: R1 no gain (OOM), G1 +1 %, N1 +3 % (within run-to-run spread: two R0 runs and T13 span 37.9–39.7). Phase 1 result: the 50 tok/s floor was not reached in the tested x8 configurations. Best measured agentic floor ~40–41 tok/s (−18 to −20 %); lab C1 55–57, lab prefill ~2,400–2,600, kit panel in band. Best config for publication: the PR #144 launch (image a21efea6, or f77c72f6 which had no OOM in its run). No T16 finalist; T18 registry recipe not prepared; the PR #144 rerun evidence and the kit submission (with this tuning table) are prepared locally for Phase 3.

0012-glm-outcome.md 0012 — GLM-5.3-Flash outcome (T24, T25a, T25b, T26)

0012 — GLM-5.3-Flash outcome (T24, T25a, T25b, T26)

Result: no GLM-5.3-Flash configuration on this box reaches the registry's 15 tok/s gate; every route is a documented experiment, and no GLM registry PR is prepared. All numbers: lab C1 decode / lab prefill, current layout.

RouteCardsDecodePrefillToolsClassification
llama.cpp b11312 IQ2_M, ub409619.6288pass, 0 parse failuresexperiment (1 card < 15)
llama.cpp IQ2_M, ub2048, GPU2-first311.8218passexperiment (3-card, < 15)
llama.cpp IQ3_XXS / IQ4_XS310.4 / 8.3132 / 156passrejected for speed, quality unmeasured
llama.cpp Q2_K / IQ3_XXS19.2 / 8.378 / 65passrejected for speed
ik_llama 5bf8f0fe IQ2_M (no MTP)313.4–14.0185fails (tool calls unparsed, 0005)experiment-only, unattested image
ik_llama IQ2_M + MTP1 / 38.2–12.6~150–166failsexperiment-only; OOM / server crash
0xSero glm53-flash-offload, exact mode (x8)17.4–7.5720–739pass, 0 parse failuresquality reference; route below gate at x8
  • T25a: no current-layout finalist (decision 0007; rolling-floor failures measured in every 10-min run).
  • T24: the uncapped candidate-vs-reference comparison runs only for finalists, so it is not run. Exact mode loads on this box (~120 GB host RAM, 23.3 GB VRAM) and stays available as the reference if a future config qualifies.
  • T25b: 1-card routes < 15 → experiment; 3-card routes < 15 → no separate recipe; ik_llama → experiment-only until a tool-parser fix and an attested local-ai-images build (operator to raise with 0xSero).
  • T26: not applicable.
  • Hardware context for the reader: dual-channel DDR4-3200 (~16.6 GB/s host memcpy measured), x8 PCIe (13.5 GB/s), 15 cores at base clock (decisions 0004, 0010). Fast mode (0xSero's 28.9 tok/s route) needs ~238 GB RAM: not possible on AM4.
Budget and stop-rule log (budget.md)

Budget

PhaseStartedBudgetStatus
0 setup2026-10-021 dayin progress
1 Flash-Next–2–3 days–
2 GLM–2–3 days–
3 review & PRs–operator-paced–
Stop-rule log
DateRouteLast 3 gainsDecision

Site rules

Status badgenot started = no records; done = a config with a kit / registry / kit+registry / experiment-only eligibility whose stability verdict is qualified; otherwise running. site/status.json overrides it (shown as override).
Stability verdictfailed = any stability run failed or crashed, or shows parse failures, repetition hits or crash/OOM; qualified = at least 3 passing runs, one of them cold, and at least 100 tool calls across them; screened = passing runs short of that.
Flash-Next targetone passing lab record of the config with decode >= 50 tok/s and lab prefill > 1000 tok/s. Headline = the qualified target-meeting config with the largest occupied context reached in its passing stability runs.
GLM targetbest lab C1 decode of a 1-card config against the registry speed gate (pins.json). 3-card configs are shown separately and do not satisfy the 1x RTX 3090 request.
Agentic-usablequalified and a passing stability run reached >= 128k occupied tokens; otherwise the label is 'configured for N k'.
Kit bandtop-1 >= 0.987 and mean KL <= 0.0013 (Flash-Next, quality_panel records).
GLM quality floorsolve-rate ratio >= 80% of the reference (metric solve_rate_ratio on quality_eval records, else computed from the model's eligibility=reference quality_eval).
Dropped configa config with a failed or crashed record and no record carrying a positive eligibility.

Pins

0xSero/local-ai-registryd21258dd744e7c78be28177c90060c6af8e10b7e
0xSero/local-ai-recipe-kitef883d269e50ecbea290f919f095e2f3ca633b42
SWE-agent/mini-swe-agent{"commit": "04d809ceab9df28f9adaed044884180159172930", "nearest_tag": "v2.4.6"}
panel qwen3.8-flash-next-exl3{"path": "reference/qwen3.8-flash-next-exl3-ref-panel.json", "sha256": "3350eef4e980779256a0c1e872d9345d0d21bdf43690e1958526aef17e02488d"}
panel glm-5.3-flash-exl3{"source": "0xSero/local-ai-recipe-kit PR #3 @ b5d0d3c2d567c4d4c1a295acc0a0fc86470b9097", "path": "artifacts/glm-5.3-flash-exl3-ref-panel.json", "sha256": "3445baefa35851f7f431b4f46fc2266a9d234d79526b1d4fc1216093f552f269"}
lab gatesload, chat, reasoning, tools, context, speed
lab speed_min_tps15.0
lab speedC1 streamed decode over first 30 s after first token, temperature 0.8, uncapped answer
lab contextneedle in a prompt of ~0.85*ctx; pass if recalled and prompt_tokens >= 0.6*ctx
lab chatfinish_reason=stop with non-empty content
lab reasoningseparate reasoning_content and 17*23=391 in content
lab toolsget_weather(city=Paris) call, then tool result (17) used in reply
lab prefill_definitioncontext-gate prompt_tokens / total request seconds (includes generation)
kit early_exit_threshold0.95
kit early_exit_policydisabled during exploration (spec §4)

pins.json sha256 858564b6e07125ac48c2ef6c5178d0a53596f21eb5f0ebda43906472b01b4d78, pinned 2026-10-02

Protocol (from SPEC.md)

2. Hardware ("the ryzen", dedicated to this work)

PartDetail
CPURyzen 9 5950X, 16C/32T, AVX2, no AVX-512
BoardASUS Pro WS X570-ACE
RAM128 GB DDR4-3200, dual-channel (4×32), ~125 GiB MemTotal
GPU3× RTX 3090 24 GB, driver 610.57.04
PCIe todayGPU0 04:00.0 x8 behind the X570 chipset; GPU1 0b:00.0 x8 CPU (drives display); GPU2 0c:00.0 x8 CPU
PCIe x16 layoutone card alone in the primary CPU slot (operator moves cards)
Storage1 TB NVMe, ~748 GB free: active models. 20 TB ZFS HDD: unused

Identify GPUs by UUID in every result. Never benchmark on the display GPU. 3-card runs are headless: the operator logs out of Hyprland and everything is driven over SSH; each record notes display: off.

4. Definitions and metrics

  • Official speed: the numbers from lab/lab.py try --on endpoint in the registry. Decode = C1 over the first 30 s at T=0.8. Lab prefill = context-gate prompt tokens ÷ total request seconds.
  • Lab prefill is a conservative number. A value > 1000 counts as proof that prefill clears 1000. Record time-to-first-token prefill separately, with the prefix cache disabled and fresh prompts.
  • Pinned harness: at Phase 0 start, record the commit SHAs of the registry, kit, mini-swe-agent and the reference panels. Every run cites them, and the six lab gates and thresholds are those at that commit.
  • Decode floor measurement: generation-only tok/s (output tokens ÷ time spent streaming them, excluding prefill, tool execution and idle time) over rolling 60 s windows, ignoring the first 60 s. "Holds the floor" = every window ≥ the floor. This is separate from official lab C1 and from the whole-run average.
  • Supplementary speed: the kit PROTOCOL.md sweep (prefill 8k/16k/32k/64k, decode C1–C4), always run with early exit disabled while exploring.
  • Real-world speed: decode measured across the whole agentic run.
  • Context is reported as occupied tokens (the prompt actually sent), alongside the configured limit. Maximum tested: 260k. The label "agentic-usable" is awarded only after a config passes §5 qualification at ≥ 128k occupied tokens; otherwise it shows "configured for N k". Shorter profiles are still published.
  • Occupied context = prompt tokens counted by the served model's own tokenizer (usage.prompt_tokens from the server), excluding generated tokens. A max-context claim needs evidence that a request actually reached that occupancy, with ≥ 4k tokens of generation headroom left.
  • MTP / speculative decoding: always report accepted tokens/s, acceptance rate and the draft settings.

5. Stability and quality protocol

Run types
  • Stability runs: capped at 10 min each (screening and qualification below).
  • Quality evaluation (GLM solve-rate, §5 GLM quality floor): a separate run with no 10-min cap but a fixed token budget per task. Finalists only.
Agentic harness
  • mini-swe-agent, pinned commit.
  • Fixed set of 5–10 SWE-bench-Lite tasks, at least one growing past 64k context.
  • Sampling is fixed: temperature, top-p, seed where supported, max tokens.
Screening run
  • One run of ≤ 10 min.
Finalist qualification
  • 3 runs of ≤ 10 min each: one starting cold, one at the claimed maximum occupied context, one repeat.
  • Cold = container restarted (prefix/KV and expert caches empty), the n-gram row cache empty, and the OS page cache dropped (sync; echo 3 > /proc/sys/vm/drop_caches). Each record lists which caches were reset.
  • Each run works through the task set in fixed order, looping, until the 10 min wall clock runs out. An unfinished task counts as unsolved.
Pass criteria (every run)
  • No crash and no OOM.
  • 0 tool-call parse failures. The ≥ 100 tool-call minimum is counted across all runs of a config (screening + qualification + extra 10-min runs as needed), not per run.
  • No repetition loops. Detector (frozen per decision 0003): any 64-token sequence repeated ≥ 3 times within one field of a response (reasoning, content, or one tool call's arguments, each checked independently; cross-field matches are logged as diagnostics only), or the same tool call (name + arguments) issued ≥ 4 times in a row within one agent conversation (the streak resets only when a fresh conversation starts). When it triggers, save the offending response. Tokens come from the served model's tokenizer when the server exposes one, else tiktoken cl100k (recorded).
  • Error categories are reported separately: malformed tool calls (invalid argument JSON, unparsed tool-call markup, unknown tool, missing command) are tool-call parse failures; responses cut off at max_tokens without a tool call are output truncations; other agent-protocol errors are agent format errors. Only the first category is the "0 parse failures" criterion; the other two are published next to it.
  • Decode stays at or above the absolute floor for the whole run: Flash-Next ≥ 50 tok/s; GLM ≥ 15 tok/s when the config is a registry candidate.
  • All six registry lab gates still pass after the run.
GLM quality floor
  • Passes the registry tools gate.
  • Meets the parse-failure rule above.
  • Solves ≥ 80 % as many tasks as the reference on the same task set, measured in the uncapped quality evaluation with identical task IDs, sampling settings and per-task token budget for candidate and reference. If the reference solves 0 tasks, the solve-rate criterion is reported as not measurable and the config cannot be a registry candidate on quality grounds.
  • Reference: GLM exact mode (bit-exact EXL3 3.05 bpw). If exact mode cannot load, the reference is the highest-bit GGUF that fits.
  • Also report KLD and top-1 agreement where they are available.
Flash-Next quality
  • Kit band: top-1 ≥ 0.987 and KL ≤ 0.0013 against the exllamav3 reference panel (tools/score_ref_panel.py, the kit's pinned panel).
  • The band is required for the headline result and for the registry PR. Configs outside it are still submitted to the kit and published, labelled "outside band" (precedent: kit PR #1).