On September 24, 2026, Perplexity announced the Fast Search option for its Search API and described Photon, the internal retrieval-and-ranking engine it says powers its production search. Perplexity reports a p99 latency of about 65 milliseconds for Photon’s stages, down from about 800 milliseconds in its previous system. Fast Search is a separate API mode built on Photon, with its own latency and benchmark figures.
What Photon does in Perplexity Search
Photon retrieves candidate web pages and ranks them for Perplexity’s search stack; it is not a standalone consumer search engine. Perplexity says the system replaced an adapted open-source engine as its index and workloads grew, and places Photon within its broader move to in-house Rust-based search infrastructure.
How a query moves through Photon
A request passes from a load balancer to Photon’s broker, which selects a group of shards. The shards retrieve and rank candidate pages; the broker then merges those candidates and fetches key fields for selected documents before returning results to the higher-level search system.
Photon uses an inverted index, which maps terms to the documents that contain them. It stores compact per-document ranking records called “docblobs” and uses different representations for short and long posting lists. A WAND-like traversal uses score bounds to limit which candidates need further ranking.
How Perplexity builds and serves its index
Perplexity separates index construction from live query serving. Pillar prepares source data in YTsaurus tables, and indexers build shard structures from that data. Versioned index files are stored for deployment; a controller rolls new index versions out across serving groups.
That separation also shapes the reported operating figures: Perplexity says Photon uses about 20% fewer equivalent serving machines than the previous system and stores about 2.5 times as much data per document. The company says it can build its full web index in a single-digit number of hours.
What Photon’s reported p99 latency measures
Perplexity’s production p99 comparison covers the retrieval-and-ranking system: about 800 ms for the previous engine and about 65 ms for Photon. The 65 ms figure covers Photon’s internal stages and excludes later stages in the higher-level search stack.
| System | Reported production p99 | Measurement scope |
| Previous retrieval-and-ranking system | About 800 ms | Retrieval and ranking in Perplexity’s production context |
| Photon | About 65 ms | Photon’s internal stages; later search-stack stages are excluded |
Fast Search benchmarks and API access
Fast Search combines Photon with a ranking configuration tuned for agent workflows. Perplexity reports 160 ms p50 and 230 ms p95 for a single Fast Search API call. Those figures describe API-call latency, not Photon’s production p99. The p50 is the median; p95 is the 95th-percentile latency.
Across six benchmarks and 3,554 selected tasks, Perplexity reports a 64.3% task score for Fast Search and 64.0% for its default preset. The company’s estimated model-plus-search costs for that benchmark set were $59.73 and $187.60, respectively. These are estimated costs across the tasks, not per-request API prices.
| Perplexity-reported measure | Fast Search | Default preset | Context |
| Task score | 64.3% | 64.0% | Six benchmarks; 3,554 selected tasks |
| Estimated model-plus-search cost | $59.73 | $187.60 | Same benchmark task set |
| Long-tail DCG relevance | 2.21 | 2.45 | Separate internal tests |
| Answer availability | 0.567 | 0.596 | Separate internal tests |
The long-tail results put the small benchmark-score difference in context: Perplexity reports lower DCG relevance and answer availability for Fast Search than for its default preset in separate internal tests.
To call Fast Search, the Search API uses search_type: "fast" on POST /search; the max_results setting can range from 1 to 20. Perplexity’s documented tariff is $1 per 1,000 successful requests.