Commit History

fix(rope): YaRN attention_factor is (0.1*ln(factor)+1)*attn_factor, not the bare multiplier
d3cca23
verified

joerowell commited on

Use chat_template.jinja as the single source: drop the {% include %} chat_template field from tokenizer_config.json
ad2b974
verified

joerowell commited on

small self-contained fixes
578e5d9
verified

joerowell commited on

Document SGLang support (sgl-project/sglang#24204)
1a001ef
verified

joerowell commited on

Note FP8 KV cache needs vLLM 0.22.0; drop scrambled-output workaround (vllm#42650)
571346d

joerowell commited on

adjust expert scales for non-Hopper targets
3ed234d
verified

joerowell commited on

update sampling parameters to match evals
d9c13a0

joerowell commited on

Update model card for 256K context length
12a5f20
verified

varunrandery commited on

increase context length to 256k
d2a5f9c

joerowell commited on

Drop VLLM_USE_DEEP_GEMM=0 from vllm serve recipe (DeepGEMM is supported on Hopper and datacenter Blackwell)
514daf4
verified

joerowell commited on

Enable thinking by default in non-Hopper FP8-KV serve command
62a5860
verified

joerowell commited on

Update non-Hopper FP8-KV serve command and link to vLLM recipes page
92f8b44
verified

joerowell commited on

Sync chat template with v5_1 (matches base/FP8/INT4)
c6a0e4c
verified

joerowell commited on

Laguna XS.2 upload
98ebde9

joerowell commited on