Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
47 stories · all sources
Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
The Open ASR Leaderboard Adds Its First Global South Language
Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
Granite 4.2 LLMs: How They're Built
Extremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC
Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
Wire It, Run It, Deploy It: AI Workflows in Gradio
Measuring benchmark optimization in speech recognition
Up to 3.2x Faster Inference with LFM2.5-DSpark
LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation
How Much Memory Does Your Agent Actually Need?
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers