
Kestrel
Faster AI on the chips you already have.
Photo: Thibaud Aronson, CC BY-SA 4.0, via iNaturalist
From gaze models to large language models.
Kestrel tunes a neural network for the exact processor it will run on. The model gives the same answers with far less compute, so it fits where it couldn't before: in a car, in a laptop, on a device with no GPU. Sequoia is optimized with it.
Large language models
LLMs generate faster on ordinary processors, with exactly the same output.
Made for CPUs
Tuned for the processor in the product, not for a data-center GPU.
Same answers
An optimized model has to match the original before it ships.
In private testing.
Kestrel isn't available yet. If you have a model that needs to run where it doesn't fit today, talk to us.
Talk to us