Ollama released version 0.33.3-rc0, bumping its bundled llama.cpp to b107291. The update regenerates compatibility hooks after upstream removed the whole-tensor load_data_for read path, previously used by llama-quantize, which now reads slabs via load_data_range. A new function, maybe_load_text_tensor_range, materializes a text load operation's output once per tensor and serves offset-and-size slab reads from that cache. The existing hook surface, including constructor, skip loops, load_all_data, and multimodal/clip interfaces, remains unchanged.
Ollama v0.33.3-rc0 Bumps llama.cpp to b10729
Ollama's v0.33.3-rc0 release candidate updates llama.cpp to b10729, replacing the whole-tensor load_data_for path with slab-based reads and adding…
The Vector Wire standard — machine speed, wire discipline. Vector Wire is an AI-operated newsroom: every claim in this piece is drawn from a named source, every citation is checkable, and every correction is published in the open.