QuTwo Unveils Quantum-Inspired Compression for AI Models
September 30, 2026 -- On the path towards full quantum advantage for AI, QuTwo is today announcing first quantum-inspired compression results for today’s most widely used AI models.
Every enterprise needs the most capable AI, but few can run it on their own terms. The largest open models need expensive, specialised AI chips just to run, which pushes most organisations to rely on closed frontier models from a handful of American labs. That means sending sensitive data outside their own control, paying for every query, and building critical capabilities on technology that competitors can rent on identical terms and that geopolitical tension can put out of reach.
We remove that dependency by making the most capable open models small enough to run on infrastructure organisations own and control, at a lower cost and a faster response.
State-of-the-art compression results across model families
Model compression makes an AI model smaller, so it needs less memory and computing power to run, while keeping as much as possible of what it can do. We have built a proprietary model compression pipeline that consistently reaches 90% compression while keeping over 90% of capability on standard reasoning and knowledge benchmarks (GSM8K and MMLU). It works across large models of different sizes and architectures, from dense models to mixture-of-experts models of more than two trillion parameters. Models that were previously out of reach for most companies to host themselves can now run on hardware they already have or can own.
Compression at this level changes the economics of enterprise AI in several ways. It unlocks use cases that never shipped because the capable model was too large for existing infrastructure. It cuts the cost of running models while holding accuracy. It brings response times down to what an end user or a real-time process requires. And it makes models small enough to deploy where they could not fit before: on-premise, on a factory floor, in a vehicle, or on a device connected to nothing at all.
"AI creates competitive advantage only when you can run it and own it. Companies should be able to run the most capable models on their own terms, on infrastructure they control. Our proprietary model compression pipeline is an answer to that," says Peter Sarlin, Founder and Executive Chairman of QuTwo.
Inside the pipeline
The pipeline has two layers. An industry-standard set of factorization, pruning, healing and quantization makes the model smaller, but loses accuracy along the way. Our proprietary Quantum-Inspired Precision-Recovery Layer then recovers much of that loss, leaving significantly less error at the same size.
In plain terms, each step removes a different kind of excess. Factorization breaks large blocks of the model into smaller, cheaper pieces. Pruning removes the parts that contribute least; in mixture-of-experts models, that means the specialist sub-networks that are rarely used. Healing lets the remaining parts readjust so they work together again, and quantization stores each number in the model with far fewer bits, down to two bits in the most compressed models. The quantum-inspired layer then corrects the errors these steps introduce.
"Standard compression methods shrink a model but leave precision on the table. Our pipeline takes that precision back, which is why we can reach a tenth of the size and keep over 90% of what the model could do," says Kuan Tan, CTO of QuTwo.
Proven at scale
The results come from an extensive evaluation campaign on Finland's LUMI supercomputer. We compressed and tested models across seven model families (YOLO, Llama, Qwen, Gemma, DeepSeek, MiniMax and GLM), covering vision, dense and mixture-of-experts architectures and sizes from 70 MB to nearly 5 TB. Each model was tested at several compression levels, and every compressed model was scored against its full-size original on the same benchmarks, hardware and settings.
Across the portfolio, models compressed by 74–80% keep 99% of their quality, and models compressed by 90+% keep 91–95%. Two models show what this looks like in practice: one of the largest open models available, and one of the hardest to compress.
- Qwen3.8-2.4T is a mixture-of-experts model with 2.4 trillion parameters and, at 4.89 TB, among the largest open models available. We compressed it to 363 GB, 92.5% smaller, while it kept 94.9% of its answer quality.
- Qwen2.5-72B is a dense model, where every part of the model is used for every answer, which makes it much harder to compress without losing quality. It was compressed from 145 GB to 14 GB, i.e. 10.4 times smaller, keeping 90.5% of its knowledge score and 93.5% of its reasoning score.
Reaching the same bar on several architectures shows the results generalize to many model architectures and categories. With our compression pipeline, the most capable AI models become something enterprises can own, run and control on their own terms, today and through the quantum transition.


