Today, Infinity Intelligence finished SFT-training the new Xora 4 Flash model. Our next-generation closed-weight model engineered for ultra-low latency, long-context reasoning and secure local deployments.
As organizations move from simple conversational models to larger scale agentic workflows, traditional Transformer architectures hit a massive scaling wall. Standard key-value (KV) caches consume quadratic amounts of VRAM, causing extreme memory requirements as the context window grows.
Xora 4 Flash breaks this bottleneck by combining a Spare Mixture-of-Experts (MoE) architecture with a linear attention mechanism inspired by RWKV. Xora 4 Flash maintains lightning-fast responsiveness across the massive 256,000-token context window, all while running completely on-premise within your own network.
Technical Highlights: Speed, Scale & Security
- Linear-Time Attention Scaling: Traditional $O(N^2)$ self-attention models become exponentially slower (and costlier) as prompt size increases. Xora 4 Flash replaces quadratic self-attention with recursive, linear-time sequence processing inspired by RWKV. Token generation remains consistent and predictable, whether processing a 2000-token problem or a 200,000-token repository.
- Efficient Mixture-of-Experts (MoE) Architecture: Dynamically routing queries through specialized sub-networks each token gives the model frontier-level knowledge at a fraction of the compute.
- Natively Multimodal & High-Precision Tool Calling: Engineered for agentic execution, Xora 4 Flash excels at tool calling, structured output generation and multi-step tool integration. The model is natively multimodal, allowing diagrams, spreadsheets and schematics to be passed directly to the model. (Full benchmark scores coming soon.)
- Built for Edge Deployments: No centralized cloud API. Xora 4 Flash is closed-weight only and always hosted locally:
- Sensitive datasets, source code and telemetry never leave your network.
- Scale up or down whenever. Local hardware means you're in control.
Hardware & Infrastructure
Xora 4 Flash is designed to run with high throughput on standard enterprise setups.
| Deployment Tier | Hardware | Quantization |
|---|---|---|
| Development & Edge | 24GB - 48GB VRAM GPU (e.g. Intel Arc Pro B70) | Q2 / Q4_K_M |
| Enterprise Production | 48GB - 96GB VRAM GPU (e.g. RTX 6000 Ada, A100) | INT8 / FP16 (BF16) |
| Scale Deployment | Multi-GPU Cluster (>96GB) | Uncompressed FP16 |
Flexible Economics on Local Infrastructure
Because Xora 4 Flash runs locally, operational costs and throughput are tied to your existing compute infrastructure.
Xora 4 Flash is currently entering its closed developer preview phase. Rather than offering off-the-shelf API keys, Infinity Intelligence works directly with technical teams to optimize deployment, infrastructure configuration and model orchestration for each specific environment.
Infinity Intelligence