HotTea LiveVerified, material updates onlyUpdated Sep 3, 4:13 PM PDT

Live

H Company releases NeoMME models for visual document retrieval

H Company's new 260M and 800M encoders use one Transformer to process text tokens and image patches. The company released the checkpoints under Apache 2.0 and added NeoMME to Hugging Face Transformers. All performance figures come from H Company.

First published Sep 3, 4:13 PM PDT · Last updated Sep 3, 4:13 PM PDT

What happened

H Company released 260M and 800M NeoMME encoders for multilingual text and images. One bidirectional Transformer processes both text tokens and image patches. The release includes pretrained backbones and visual document retrieval checkpoints. H Company also added NeoMME to Hugging Face Transformers.

Why it matters now

The 260M model gives document-search teams a smaller open model for visual retrieval. H Company reports that it processes 51 pages per second on one NVIDIA L40S at 2048 by 2048 resolution. It also reports cutting late-interaction storage from about 1.5 MB to 6 kB per page. The company says the smaller files retain more than 95% of its baseline retrieval score. The authors ran these tests, and the checked sources include no independent reproduction.

Updates

What changed

H Company releases NeoMME models for visual document retrieval

H Company released 260M and 800M NeoMME encoders and visual retrieval checkpoints. The Apache 2.0 models use one Transformer to process text tokens and image patches. H Company also added NeoMME to Hugging Face Transformers.

Verification

Primary evidence before publication.

Social chatter can identify a lead. It does not authorize a HotTea live story.