oss

libvips keeps a large image from becoming a stack of large images

8 sources 7 primary sources September 25, 2026

Loading reads and saves…
Text
Archival black-and-white photograph of the VASARI camera-positioning gantry, rails, lighting equipment, and cables at the National Gallery.

The VASARI positioning equipment and lighting system, reproduced from figure 1, page 72, of David Saunders and John Cupitt's 1993 National Gallery Technical Bulletin article. This is the imaging project for which VIPS began.[1][2]

A photograph can become expensive long before anyone sees it. Decode it, resize it, adjust it, and save it: if each stage retains a complete pixel array, the working set contains several versions of the same picture. The open-source library libvips changes that arrangement. It connects operations so that small regions can travel through the calculation together. Understanding which regions get requested—and who requests them—explains both its economy and its limits.[1][3]

The problem once stood in front of a painting. VIPS began in 1989 for VASARI, a project making high-resolution, multispectral images of artworks. The National Gallery's 1993 account describes a camera on positioning equipment, recording overlapping portions that software joined into a complete image. The photograph above shows that equipment: rails, cables, and a gantry built to measure a surface carefully.[1][2]

That origin gives the architecture a tangible purpose. A painting could yield more image data than a workstation could comfortably hold. The following explanation uses the project's 8.17 documentation and its Python binding examples to trace how a large logical image can exist without every intermediate pixel occupying memory at once.

An image can hold a recipe

In libvips, a region is a rectangle of pixels. A partial image can hold the functions needed to produce such rectangles, instead of a complete pixel array. In the C interface, vips_region_prepare() asks for a particular rectangle; the library arranges to make those pixels available.[3]

Suppose an output region needs a color adjustment. That operation requests the corresponding input region. If its input is another calculated image, the request continues upstream. The chain describes complete images, but the live intermediate storage can consist of the small areas currently being processed.

Evaluation needs a consumer, called a sink. Writing a file is one sink; displaying pixels, producing a memory image, or calculating statistics can also pull work through the pipeline. Lazy evaluation therefore does not mean that only saving can trigger computation. It means that something must demand results.[3]

The last line pulls on the earlier ones

This small Python example uses the documented pyvips loading, arithmetic, and file-output interfaces. Assume an ordinary three-band RGB JPEG:

import pyvips

source = pyvips.Image.new_from_file("painting.jpg", access="sequential")
adjusted = source * [1.0, 0.9, 1.0]
adjusted.write_to_file("painting-adjusted.jpg")

The multiplication reduces the green channel. It is a deliberately simple demonstration of dataflow, not a conservation-grade color correction. Opening the file reads its header; the arithmetic connects another operation; writing the output demands the resulting pixels. The filename suffix selects the output format.[4]

A variable such as adjusted consequently need not mean “a second decoded photograph is sitting in RAM.” It can describe how to obtain that photograph. This distinction matters when reading application code: a sequence of apparently whole-image operations does not necessarily imply a sequence of whole-image allocations.

The threading follows the same organization. The project paper describes workers carrying lightweight copies of the pipeline's writable state. Different output regions can be calculated through the chain concurrently, without assigning each processing stage a separate full-size image. This is one reason region size and pipeline structure matter alongside the number of CPU cores.[1]

The file format still gets a vote

The loader determines what requests the source can satisfy. A tiled TIFF can provide individual tiles through its loading library. Other formats or loaders may require full decompression into a representation that supports random access. Depending on size and configuration, that representation can live in RAM or a temporary disk file.[5]

The access="sequential" hint tells the loader that the calculation will move through the image from top to bottom. Where supported, this lets decoding feed a pipeline incrementally. It is a statement about the program's access pattern, not a command that makes every codec stream.[4][5]

A vertical flip exposes the constraint. To write the result from top to bottom, the program needs the source from bottom to top. The JPEG decoder described in the documentation reads forward. The system therefore needs a random-access representation before it can complete that transformation.[5]

The practical consequence is easy to miss: two operations on the same compressed file can have different storage requirements. A tiny source file also says little about the size of its decoded pixels. When temporary disk usage rises, the explanation may be the requested access order rather than an unexpectedly large intermediate operation.

The destination can ask for everything

The output contract can restore a whole-image memory requirement. write_to_memory() produces an unformatted pixel array; write_to_buffer() instead produces an encoded image in a memory buffer. Those are different allocations, and neither is the same as streaming a file to disk.[6]

For scale, a hypothetical 20,000 × 20,000 image with three eight-bit channels needs 1.2 billion bytes for its raw pixels alone: width × height × channels × bytes per channel. That is arithmetic, not a measured libvips benchmark. Avoiding large intermediates cannot remove a final array that the caller explicitly requests.

Cache controls address another part of the working set. vips_cache_set_max_mem() sets a threshold at which cached operations start being dropped. The documentation explicitly excludes memory allocated by external libraries from that accounting. Treating the setting as a hard ceiling on process memory therefore misreads its scope.[7]

A benchmark needs the whole trip

An independent engineering account from the Studysapuri product team helps connect this design to an application. In March 2021, Yuta Hamada described replacing a Go imaging library with bimg, a libvips binding, for thumbnail generation. The comparison used specified inputs and 200-pixel-wide outputs, with JPEG quality explicitly set on both implementations.[8]

The account reported lower processing time and memory use for its tested workload. It also examined decoder behavior, concurrency, and operational details. That is useful external evidence, with a defined task attached; it does not establish a universal speed multiplier for every image transformation.[8]

For a small team maintaining an image service, the corresponding evaluation can stay concrete: test representative formats, transformations, and destinations, then observe process memory and temporary storage under realistic simultaneous requests. Include the awkward inputs that change access order. A team without visibility into those resources cannot validate its memory budget from the cache setting alone.

The VASARI photograph captures a machine moving carefully across a painting. Its software descendant asks similarly bounded questions of an image: which rectangle is needed next, what must be read to produce it, and where will the result go? Following those questions all the way from decoder to destination makes the memory behavior intelligible.

Sources

  1. John Cupitt, Kirk Martinez, Lovell Fuller, and Kleis Wolthuizen, “The libvips Image Processing Library,” Electronic Imaging, 2025 — project origins, demand-driven processing, and horizontal threading.
  2. David Saunders and John Cupitt, “Image Processing at the National Gallery: The VASARI Project,” National Gallery Technical Bulletin 14, 1993, pp. 72–85 — camera acquisition system and archival equipment photograph, figure 1 on p. 72.
  3. libvips 8.17, “Technical background: Evaluation” — regions, partial images, operation pipelines, and sinks.
  4. pyvips, “Introduction” — lazy loading, sequential access, arithmetic on bands, and writing images.
  5. libvips 8.17, “Technical background: Opening files” — tiled access, full decompression, temporary storage, and the vertical-flip example.
  6. libvips 8.17, output API documentation — raw pixel arrays with write_to_memory() and encoded image buffers with write_to_buffer().
  7. libvips 8.17, cache_set_max_mem() — operation-cache eviction threshold and the exclusion of external-library allocations.
  8. Yuta Hamada, “Speeding up thumbnail creation with bimg (libvips Go bindings),” Studysapuri Product Team Blog, March 1, 2021 — independent implementation account and explanation of libvips; Japanese-language source.
Previous Therion lets a cave map change without starting over Next LAMMPS divides the box, then shares the neighbors

Recommended In oss

Matched by subject and format