Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

PyTorch Consolidates Media Stack Into TorchCodec, TorchVision, TorchAudio

PyTorch has consolidated media decoding and encoding into TorchCodec, narrowing TorchVision and TorchAudio to transforms, with all three libraries now…

The PyTorch team said it has consolidated its media processing capabilities across three libraries: TorchCodec for all decoding and encoding of images, video, and audio on CPU and CUDA; TorchVision for image and video transforms; and TorchAudio for audio transforms1. All three libraries are now ABI stable, and previous decoding and encoding APIs in TorchVision and TorchAudio have been deprecated or removed. TorchCodec is described as generally more performant than prior implementations, particularly for CUDA video decoding. The remaining functionality in TorchVision and TorchAudio outside transforms is no longer under active development.