A GitHub gist demonstrating how to run Llama (OS Jev) offline on a Mac with M4 chip using Apple's CoreML framework, achieving 45 tokens per second inference speed. The project showcases local LLM deployment leveraging Apple's neural engine for efficient on-device inference.
Background
Apple's CoreML framework enables developers to run machine learning models locally on iOS and macOS devices using the Neural Engine. Running large language models offline on consumer hardware is an active area of interest as privacy and latency concerns grow.
- Source
- Hacker News (RSS)
- Published
- Sep 20, 2026 at 11:58 PM
- Score
- 6.0 / 10