https://llamabench.ai/browse/rtx4090/qwen3-8-27bConsumer GPU Hardware Benchmarks | llm-bench.io sergiuszm/ninfer-4090: Qwen3.8-27B on one RTX 4090: full native 262K context via E8 4-bit KV, up to 149 tok/s code decode with MTP3, sm_89-retuned attention prefill…