running Underdog-Saluki-27B on a single RTX 3060. 12GB VRAM. (config below)
it’s Qwen3.8-27b squeezed into 7.9GB (2-bit) and tuned to keep tool calling intact. dense, so every layer sits on the gpu. no expert offload.
170K context. ~22 tok/s decode, ~490 tok/s prefill.
running my agentic coding benchmark on it now. will compare it against other Qwen3.8-27B quants next!
config:
llama-server
-m Underdog-Saluki-27B-1.0-IQ2-mix.gguf
-ngl 99 -c 170000 -fa on --jinja
-np 1
--cache-type-k q4_0 --cache-type-v q4_0
-b 2048 -ub 512
--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0
--repeat-penalty 1.0
--reasoning on
--chat-template-kwargs ‘{“reasoning_effort”:”medium”}’
@UnderdogAI
it’s Qwen3.8-27b squeezed into 7.9GB (2-bit) and tuned to keep tool calling intact. dense, so every layer sits on the gpu. no expert offload.
170K context. ~22 tok/s decode, ~490 tok/s prefill.
running my agentic coding benchmark on it now. will compare it against other Qwen3.8-27B quants next!
config:
llama-server
-m Underdog-Saluki-27B-1.0-IQ2-mix.gguf
-ngl 99 -c 170000 -fa on --jinja
-np 1
--cache-type-k q4_0 --cache-type-v q4_0
-b 2048 -ub 512
--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0
--repeat-penalty 1.0
--reasoning on
--chat-template-kwargs ‘{“reasoning_effort”:”medium”}’
@UnderdogAI
8 5 0 59 5.7K 52
Underdog AI