Perplexity Released Photon Retrieval Engine

The proprietary engine reduces retrieval latency and infrastructure costs for AI search queries.

Updated on Sept. 30, 2026 in Artificial Intelligence

Isometric editorial illustration showing interconnected server blocks and storage nodes, representing an enterprise search retrieval engine.
Perplexity has deployed its new Photon retrieval engine across its production environment, aiming to resolve latency issues and optimize operational scaling for AI-driven search tasks. AI Illustration. Upload story photo >

Live Poll

Would you trade search result accuracy for faster and cheaper AI search tools?

Perplexity has released Photon, a proprietary Rust-based retrieval and ranking engine that now handles all production traffic. The system replaces a previously forked open-source engine to improve performance and operational efficiency.

Why it matters

The transition addresses critical tail latency and merge spikes that hampered the previous architecture. By optimizing the retrieval stack, the firm seeks to sustain performance while scaling agent-based search tasks.

Photon cut p99 latency to 65 ms from 800 ms and reduced the required server count by 20%. While the engine cut agent-task costs by approximately 68%, internal long-tail benchmarks showed a relevance drop from 2.45 to 2.21.

The players

Perplexity

An AI company focused on conversational search engines that integrate real-time web indexing with large language models.

The details

The engine uses a load balancer to route queries to a broker, which fans out requests to shard groups for retrieval and ranking. Indexers build versioned shard indexes from YTsaurus tables—a distributed storage system—on dedicated nodes to manage the data. This architecture replaces the previous system, which suffered from merge spikes and slow recovery times.

Timeline

  1. September 30, 2026: Perplexity released the Photon engine.

The Tech Race

Perplexity is moving away from forked open-source foundations toward highly specialized, proprietary infrastructure. This mirrors a broader industry trend where search companies optimize custom retrieval engines to lower latency and compute overhead.

Users experience the engine via the new Fast Search mode, which is priced at $1 per 1,000 requests. This tier is optimized for speed and cost-efficiency in agent-based tasks rather than maximum long-tail relevance.

The takeaway

The move to Photon underscores the importance of low-level latency optimization in the competitive AI search space. Watch for future benchmark updates to see if the firm reconciles the drop in long-tail relevance with the new speed gains.

Further reading

For broader context on how search architectures are evolving, visit Artificial Intelligence.

Source note: This article includes information reported by MarkTechPost.

Live Poll

Would you trade search result accuracy for faster and cheaper AI search tools?