Files
mlx-lm/ACKNOWLEDGMENTS.md
5fa62eb5f5 Add bitnet1.58 with custom metal kernel (#219)
* add bitnet

* update activation to relu2

* working bitnet

* remove artifacts

* remove logging

* add custom post quant

* fix dtype and add compile

* fixed weight unpack

* add custom kernel to avoid memory overhead

* compile relu2

* fix weight scale

* remove unused

* add tests and update tuner utils

* update acknowledgements

* add kernel caching

* add act_quant and set float16 as default dtype

* use mx.add and move scaling to kernel

* remove act quant

* move bitlinear layers to separate file

* feat: add falcon-e and other bitnet support

* refactor: address comments

* add support for 1.58bit N-bit quants

* 43.85% speedup in generation performance (M3 max)

* refactor utils

* remove masking (2% gen speed improvement)

* add quantization config

* test llama bitnet

* refactor apply_hf_quant

* default threadgroup: 64 -> 32

* add comment

* fix prompt processing perf

* remove modulo

* compile kernel in the constructor

* Improve the bitnet kernel

* remove benchmark

* refactor bitlinear swap

* format

* remove llama changes

* revert utils

* faster + cleanup

* not trainable

* fix tests

---------

Co-authored-by: younesbelkada <[email protected]>
Co-authored-by: Angelos Katharopoulos <[email protected]>
Co-authored-by: Awni Hannun <[email protected]>
2025-07-02 16:34:07 -07:00

1.1 KiB

Individual Contributors

If you wish to be acknowledged for your contributions, please list your name with a short description of your contribution(s) below. For example:

  • Jane Smith: Added the foo example.

MLX LM was developed with contributions from the following individuals:

  • Shunta Saito: Added support for PLaMo models.
  • Gökdeniz Gülmez: Added support for the following architectures: OpenBMB's MiniCPM and MiniCPM3, Kyutai's Helium, State-Space'sMamba v1, Z.ai & THUKEG's GLM4, Rednote dots.llm1, and Allenai's OLMoE; Added support for the following training algorithms: full-fine-tuning; Added support for the following other features: Multiple Optimizers to choose for training, and reporting training metrics to WandB (Weights & Biases).
  • Prince Canuma: Helped add support for the following model architectures: HuggingFace's Starcoder2, Cohere's Cohere (1 and 2), Alibaba Qwen's Qwen (2, 3 and MoE), Microsoft's Phi (3 and 3.5 MoE), BitNet1.58, Meta's Llama (3 and 4), Google DeepMind's Gemma 3, and InterLM's InternLM 2.5.