Data Science Wire

Qwen3.8 Flash Next llama.cpp config tuning

Reddit r/LocalLLaMA3d4 min read

Hola all. Do you guys mind sharing your LLama.cpp config and system setup details for Qwen3.8 Flash Next? Model's quite big and tryining many combinations of llama.cpp options takes lots of time, so looking for other people setup details. I've attached my current config at the bottom, so if anyone sees something that could be improved please shout. My current best result: - PP within 130...200 tps (limited by cpu?) - TG within 14..22 tps (~15tps on average) Hardware: - Dual RTX 3090 (48GB VRAM) - 128GB DDR4 - Some old Xeon 40 core - Proxmox VM, pcie passthrough, numa binding to a single phys c

Read the full story at Reddit r/LocalLLaMA

More in AI