Community MLX runtime for local Laya inference. Fast short decisions on a Mac. Not an official Convai build.
What Laya-MLX adds
The upstream Laya stack typically uses PyTorch and Transformers. Laya-MLX ports weights and the decision heads to MLX so developers on Mac can run inference without a CUDA box. Tokenization still uses Hugging Face tokenizers.
The project documents numerical parity checks against upstream on fixed fixtures.
Published latency (M3 Max)
md in the laya-mlx repo measures end-to-end short-input latency on Apple M3 Max (40-core GPU, 128 GiB RAM). 4 ms (laya-multilingual), with model load excluded. Ten-question batches and full-context runs are much slower. These numbers are one machine and one OS build.
They are not a guarantee on M2, M4, or laptops with less memory.
Comparison to PyTorch MPS on the same Mac
The same document shows PyTorch MPS FP32 somewhat slower than MLX FP16 on identical hashes. That is expected when FP16 is allowed. Throughput rows use repeated question templates. Distinct-question production mixes may differ.
Relationship to cloud Jev
Community posts (including our showcase blurb for Lonely__MH's snake demo) sometimes contrast local tens-of-ms inference with hosted Jev latency. That is apples to oranges. No TLS, no multi-tenant queue, different hardware class. Use Laya-MLX when Mac-local loops matter.
Do not claim a universal 50x speedup over your Jev region without measuring.
Getting started
Clone github.com/mizorewww/laya-mlx, follow uv sync extras in the README, and pin upstream Laya revisions the benchmark scripts expect. PyPI lists laya-mlx for packaged installs. For Convai's official path on Linux/GPU, use the laya PyPI package instead.
Showcase: snake on MLX
jev.aitools.fyi embeds a public X clip where snake ticks use local Laya decisions. Useful for sensing decision frequency, not for SLA proof: /showcase?demo=laya-mlx-lonely-mh.
More in this hub
Related Jev learn guides
Live demo
Lonely__MH's laya-mlx snake clip shows local typed decisions driving a game loop on Apple Silicon. Not a benchmark, but a useful feel for throughput.
FAQ
- Is Laya-MLX official?
- No. Convai ships the PyTorch/Hugging Face stack. Laya-MLX is an independent MLX port maintained separately.
- Do I need PyTorch to run Laya-MLX?
- Inference is MLX-first. The repo uses PyTorch MPS as a reference backend in benchmarks, not as a production requirement for MLX mode.
- Will MLX match T4 GPU numbers on the HF card?
- Different hardware. HF documents about 33 ms on Tesla T4 for routed GPU inference. MLX documents about 7-16 ms P50 short paths on M3 Max. Neither predicts your laptop under load.
Primary sources
- Hugging Face: convaiinnovations/laya (model card and benchmarks)
- GitHub: NandhaKishorM/laya (upstream SDK)
- PyPI: laya package
- TypeSafe docs: System One / Jev API
- Laya-MLX: BENCHMARKS.md (M3 Max, independent port)
- GitHub: mizorewww/laya-mlx
- PyPI: laya-mlx
- Third-party explainer: What is Laya-MLX? (aiidelist.com)