A common kestrel perched on a post against a deep blue sky, its tail fanned.

Kestrel

Faster AI on the chips you already have.

Photo: Thibaud Aronson, CC BY-SA 4.0, via iNaturalist

From gaze models to large language models.

Kestrel tunes a neural network for the exact processor it will run on. The model gives the same answers with far less compute, so it fits where it couldn't before: in a car, in a laptop, on a device with no GPU. Sequoia is optimized with it.

  • Large language models

    LLMs generate faster on ordinary processors, with exactly the same output.

  • Made for CPUs

    Tuned for the processor in the product, not for a data-center GPU.

  • Same answers

    An optimized model has to match the original before it ships.

In private testing.

Kestrel isn't available yet. If you have a model that needs to run where it doesn't fit today, talk to us.

Talk to us