Laya-CoreML
Apple Silicon Core ML and Neural Engine runtime for open-weight Laya: short typed decisions, PyPI laya-coreml, Hugging Face bundles, Snake demo.
Overview
Laya-CoreML (github.com/mizorewww/laya-coreml, Apache-2.0, PyPI laya-coreml) is an independent Core ML port of Convai Innovations Laya, written beside the mizorewww/laya-mlx sibling. pip install laya-coreml, then laya.load pulls a Hugging Face bundle such as aac6fef/laya-multilingual-coreml-ane. predict() returns probabilities for choice, ordinal score, and boolean noul. There is no autoregressive decode and no JSON to parse. The ANE FP16 bundle caps the whole request at 96 tokens, counting question, options, and state. Longer text raises a capacity error. The general multilingual CPU and GPU bundle aac6fef/laya-multilingual-coreml keeps a 1024 token window. README hardware note: M3 Max, 40-core GPU, 128 GiB, macOS 27.2. One 91-token question padded to 96, with prompt prep through formatting and with load excluded, measured 4.98 ms P50 and 5.31 ms P95 on ANE FP16 versus 6.94 ms and 7.39 ms for compiled MLX FP16, over 65,598 stable calls. Whole-system energy per decision was 0.1540 J versus 0.4288 J (2.78x), from SMC PSTR samples with anomaly rejection. A W8 palette variant reached 4.88 ms P50 and 3.19x energy. README states the requested 10x speedup was not achieved. Following upstream v0.3.5, fitted calibration temperatures are clamped to the range 0.5 through 5.0 because the shipped choice:11+ bucket is 0.1006, which would report a coin flip as near certainty. Raw buckets stay on agent.temperature_raw. Public GitHub counts were 1,527 stars and 128 forks when this listing was drafted. The project states it is not an official Convai Innovations or Apple release.
Problem: Short Laya decisions on a Mac still default to MLX or a CPU ONNX session, even when the Neural Engine is the chip you actually paid for.
Built for: Apple Silicon developers on macOS 15+ and Python 3.11 to 3.13 who want local choice, score, and noul calls from published Core ML bundles, plus a terminal Snake loop that shows the probabilities.
First indexed on Jev Directory: 2026-09-29
Creator and team
- Name
- mizorewww
- Handle
- @mizorewww
How Jev is used
- Role in the product flow
- On-device Core ML System One: one forward pass for choice, score, and noul on Apple Silicon, with an ANE path for short requests
- Primitives
- ChoiceScoreNoul
- State in
- Plain text or structured state plus a questions map (choice, score, or noul). ANE bundles enforce a 96 token total budget across question, options, and state.
- Decision out
- Probability answers per question key. No generated tokens. Calibration temperatures outside 0.5 to 5.0 are clamped, with a RuntimeWarning naming each bucket.
- pip install laya-coreml on Apple Silicon, macOS 15+, Python 3.11 to 3.13
- laya.load a Hugging Face id such as aac6fef/laya-multilingual-coreml-ane, or a local directory with local_files_only
- Call predict with state text and a questions dict for choice, score, or noul
- For the Snake loop, pip install laya-coreml[demo], hf download the snake bundle, then laya-coreml-snake
- Export your own graph with laya-coreml[convert] and laya-coreml convert when you need a custom bundle
This page is the Core ML and Neural Engine runtime. receptron-laya embeds Convai Laya ONNX inside Node. ollaya-dev-ollaya pulls many open decision graphs and serves POST /v1/systemone on port 11435. The learn guide at /learn/laya-vs-jev/laya-mlx-apple-silicon covers the MLX sibling (github.com/mizorewww/laya-mlx), including its own 7 to 16 ms short-decision notes. Laya-CoreML ships separate Hugging Face bundles: English 421M at 512 tokens, multilingual 322M at 1024, typed-decisions 421M at 1024, a Snake GPU pack, and two 96-token ANE packs (FP16 and approximate W8). Ordinary SDPA export stays on CPU plus GPU after RangeDim GPU shapes failed local fidelity checks. The ANE rewrite uses BC1L activations, 1x1 projections, and per-head attention. CPU still handles input and output boundaries. Cite mizorewww/laya-coreml README, docs/ANE_BENCHMARKS.md, and docs/USAGE.md. Upstream weights trace to Convai Innovations Laya (github.com/NandhaKishorM/laya).
Sourced performance claims
- ANE FP16 P50/P95 4.98/5.31 ms versus compiled MLX FP16 6.94/7.39 ms on one short multilingual question; system energy 0.1540 J versus 0.4288 J (2.78x). W8 palette: 4.88/5.23 ms and 3.19x energy. 65,598 stable calls on M3 Max. The 10x target was not met.Source: github.com/mizorewww/laya-coreml README Measured on M3 Max
- Three general-purpose FP16 checkpoints match upstream selected answers on 189/189 validation questions. ANE FP16 L96 passes 59/59 fitting questions with max calibrated-probability drift 0.002925. W8 drift is 0.014393 under a 0.02 gate. Six-bit and four-bit packs failed that gate and are not published.Source: github.com/mizorewww/laya-coreml README Port fidelity and limits
- Snake loop: 49.1 to 50.0 decisions/s across three uncapped 600-step episodes, zero deaths, two safety interventions. A paired 600-step check matches 600/600 actions. A 1024-token ANE request is about 91.7 ms in its serial screen.Source: github.com/mizorewww/laya-coreml README and docs/SNAKE_BENCHMARKS.md
- Public GitHub repo mizorewww/laya-coreml had 1527 stars and 128 forks when this listing was drafted.Source: GitHub API 2026-09-29
Features and stack
Features
- PyPI package laya-coreml with predict() for choice, score, and noul
- Hugging Face bundles with tokenizer, model card, checksums, and packaging-time validation
- ANE FP16 and W8 short-context packs plus CPU and GPU packs up to 1024 tokens
- Terminal Snake demo with visible probabilities, score, latency, and a cycle safety layer
- Calibration clamp to 0.5 through 5.0, with raw temperatures still readable
- Reproducible speed and energy benches under benchmarks/results
Stack
- Python 3.11 to 3.13
- Apple Core ML
- Apple Neural Engine for short ANE bundles
- Hugging Face weights under aac6fef
Pricing: Apache-2.0 code. Inference stays on your Mac. Hub downloads are the bandwidth cost. README labels the port independent of Convai Innovations and Apple.
Links
FAQ
- Is this the Laya-MLX learn page or the Node client?
- No. /learn/laya-vs-jev/laya-mlx-apple-silicon documents mizorewww/laya-mlx. receptron-laya is the npm ONNX library. ollaya-dev-ollaya is the multi-model daemon. This listing is the Core ML and Neural Engine product in mizorewww/laya-coreml.
- Why did a long prompt fail on the ANE model?
- README: the ANE bundle has a 96 token total limit, including question, options, and state. Use aac6fef/laya-multilingual-coreml for the 1024 token general model.
- Did they hit a 10x speedup versus MLX?
- README says the requested 10x improvement was not achieved. The published short-question gain is about 1.39x versus compiled MLX FP16, with about 2.78x better whole-system energy per decision on that same experiment.
- Are the Snake frames a latency certificate?
- README separates them. The 4.98 ms figure is one short question. The Snake loop includes rendering serialization and ran at 49.1 to 50.0 decisions per second in three 600-step episodes. A full-game ANE speedup over compiled MLX is not claimed.
Related learn guides
Original Jev guidance that pairs with this product pattern.
- Where to run Jev
Compare official TypeSafe, Vercel AI Gateway, OpenRouter, Cloudflare Workers AI, and classifier.dev with a fact table.
- System One model
The model family behind Jev: parallel typed questions, one forward pass, probabilities you can threshold.
Related products
Hand-picked neighbors with rich profiles or overlapping tags.