Pinned Loading
-
turboquant-plus
turboquant-plus PublicTurboQuant+: 3-bit KV cache value quantization and group size optimization for long-context LLM inference
Python 1
-
ChannelQuant
ChannelQuant PublicNear-lossless 4× KV-cache compression for GQA models at ~4.1 bits/value — per-channel-key INT4 + static outlier ROM, with a reproducible reference model and paper.
Python
-
LonghornSilicon/lambda-kve
LonghornSilicon/lambda-kve PublicRead-only mirror of the KVE (KV Cache Engine) block of the Lambda monorepo (LonghornSilicon/lambda). ChannelQuant KV-cache codec. Develop in the monorepo; this repo is auto-mirrored.
Python
-
LonghornSilicon/lambda-acu
LonghornSilicon/lambda-acu PublicRead-only mirror of the ACU (Attention Compute Unit) umbrella block of the Lambda monorepo (LonghornSilicon/lambda). MatE + VecU + precision controller — the Q·Kᵀ→softmax→P·V datapath. Develop in t…
Python
-
covenant
covenant PublicContract-driven GPU analytical placer (DREAMPlace fork overlay) with a noise-calibrated evaluation protocol, paired multi-seed CIs, liveness gates, calibration arms.
Python
If the problem persists, check the GitHub status page or contact support.