A deep dive into how Perplexity serves search results at scale: embeddings for ranking, GPU-based mo...
Aravind Srinivas(@AravSrinivas) · 人工智能
A deep dive into how Perplexity serves search results at scale: embeddings for ranking, GPU-based model inference, request batching, running inference servers, and handling latency/throughput trade-offs. Perplexity @perplexity_ai Every answer in Perplexity starts with embedding and ranking models picking the most relevant results for the query. Today we published research on how we built SoTA serving infrastructure behind those models. Read the research: perplexity.ai/hub/blog/fast-… 🔗 View Quoted Tweet 💬 8 🔄 1 ❤️ 65 👀 10115 📊 10 ⚡ Powered by xgo.ing