15 curated breaking changes across major versions of sglang. Use this as a migration checklist before bumping dependencies.
**`helion` jumps 0.2.6 to 1.4**, a major-version move for anyone depending on helion-backed kernels: [#32562](https://github.com/sgl-project/sglang/pull/32562)
**sgl-kernel AOT `bmm_fp8` is deleted in favor of `flashinfer.bmm_fp8`; the AOT router GEMM and fused A GEMM are also removed**: [#31202](https://github.com/sgl-project/sglang/pull/31202), [#30280](https://github.com/sgl-project/sglang/pull/30280)
**CuteDSL BF16 GEMM on SM100 is on by default** when the heuristic allows it: [#30567](https://github.com/sgl-project/sglang/pull/30567)
**Breakable prefill CUDA graph is now on by default for DP attention**: [#31682](https://github.com/sgl-project/sglang/pull/31682)
**`sglang.jit_kernel` is retired into `sglang.kernels`**, completing RFC [#29630](https://github.com/sgl-project/sglang/issues/29630). Imports from the old module path must move: [#32072](https://github.com/sgl-project/sglang/pull/32072), [#31666](https://github.com/sgl-project/sglang/pull/31666), [#32015](https://github.com/sgl-project/sglang/pull/32015), [#32045](https://github.com/sgl-project/sglang/pull/32045)
**Chunked input-logprob processing is now on by default** to cap peak memory: [#31498](https://github.com/sgl-project/sglang/pull/31498)
**The experimental QServe (QoQ) W4A8 and FBGEMM FP8 quantization paths are removed** (per [#28543](https://github.com/sgl-project/sglang/issues/28543)): [#31109](https://github.com/sgl-project/sglang/pull/31109)
**CUTLASS FP8 blockwise deleted for SM90 / SM100**, SM120 moved to JIT: [#30438](https://github.com/sgl-project/sglang/pull/30438)
**`--fp4-gemm-backend cutlass` is removed** along with the in-tree NVFP4 JIT kernels, so NVFP4 GEMM now requires FlashInfer. Use `auto`, which picks `flashinfer_cutedsl` on SM100 and `flashinfer_cutlass` on SM120: [#30448](https://github.com/sgl-project/sglang/pull/30448)
**`UnifiedRadixTree` is now the default** for SWA, Mamba and DSA models. A behavior change on those architectures: [#30468](https://github.com/sgl-project/sglang/pull/30468)
from sglang.srt.entrypoints.openai.protocol import Tool ```
from sglang.srt.openai_api.protocol import Tool
Update Prometheus dashboards for new metrics
Review new error response format ## ⚡ Built for speed. Engineered for scale. Production-proven.
Update worker API integrations for UUID-based management
Get this data programmatically \u2014 free, no authentication.
curl https://depscope.dev/api/breaking/pypi/sglang