LinkedIn Boosts Semantic Search with Efficient GPU Retrieval

Dhritiman Das, Chujie Zheng, Ronak Kaoshik, Pratik Dixit, Vishal Shah, Yanbo Li, Jiahao Xu, Manika Agarwal, Chinmay Naik, Lingyu Zhang, Chetan Bhole, Chirag Bhanuprasad Mehta, Meng Zheng, Puneet Singh Ahluwalia, Shirisha Singh, Ping Jin, Manas Apte, Gokulraj Mohanasundaram, Tugrul Bingol, Raghavan Muthuregunathan, Fedor Borisyuk· September 1, 2026 View original

Key takeaways

  • LinkedIn improved semantic search with a policy-aligned GPU retrieval framework.
  • The system uses category-supervised embedding segments and a two-stage GPU architecture.
  • FP8 coarse ranking and FP16 re-ranking boost capacity and throughput.
  • Significant gains in search relevance and precision were observed in live tests.

Who benefits

Social MediaRecruitingE-commerceEnterprise SearchAdTech

Summary

LinkedIn developed a policy-aligned GPU retrieval framework for semantic search, partitioning embeddings into category-supervised segments. This two-stage GPU architecture, using FP8 and FP16, significantly improves offline relevance and live A/B test precision for complex natural-language queries.

LinkedIn has developed a new, highly efficient GPU retrieval framework designed to enhance semantic search capabilities, particularly for complex natural-language queries like "a fintech founder in Berlin who worked in payments." Traditional cosine similarity often averages evidence, potentially masking a failure on one critical facet with a strong match on another, thereby limiting recall. To address this, LinkedIn's new framework aligns with a bottleneck-oriented relevance policy, where every non-negotiable facet must be satisfied. The core of this framework involves partitioning embeddings into eight category-supervised segments. Scores for these segments follow a min/median rule during serving, and for multi-vector retrieval, the segment score is computed independently per tagged document slot and maximized across slots. A lightweight single-slot Stage-1 scorer generates high-recall candidates, while scale-invariant relative-norm gating ensures consistent category activation. This framework is served using a two-stage GPU architecture. An FP8 coarse ranker efficiently scores the entire corpus, boosting per-shard capacity by 71% and Stage-1 matmul throughput by 36%. Subsequently, an FP16 stage precisely re-ranks an oversampled candidate set, recovering nearly all full-FP16 recall at over 500 queries per second per shard replica. Offline evaluations on 21,000 held-out queries showed improved relevance, and a live A/B test demonstrated significant gains in exploratory-query Precision@10 and navigational Precision@1, confirmed by human evaluation.

Why it matters

This advancement provides a blueprint for other large-scale platforms to implement highly efficient and accurate semantic search, directly improving user experience and the effectiveness of talent discovery or product matching.

How to implement this in your domain

  1. 1Adopt multi-stage retrieval: Investigate implementing a two-stage retrieval architecture (coarse and fine ranking) to balance efficiency and accuracy for large-scale search.
  2. 2Leverage GPU acceleration: Explore using GPU-optimized retrieval for embedding-based search to handle massive corpora and high query loads.
  3. 3Align retrieval with policy: Design embedding and scoring strategies that directly reflect business relevance policies, such as ensuring all critical facets are met.
  4. 4Experiment with mixed precision: Evaluate the use of mixed-precision (e.g., FP8 for coarse, FP16 for fine) inference to optimize throughput and resource utilization.

Original post by Dhritiman Das, Chujie Zheng, Ronak Kaoshik, Pratik Dixit, Vishal Shah, Yanbo Li, Jiahao Xu, Manika Agarwal, Chinmay Naik, Lingyu Zhang, Chetan Bhole, Chirag Bhanuprasad Mehta, Meng Zheng, Puneet Singh Ahluwalia, Shirisha Singh, Ping Jin, Manas Apte, Gokulraj Mohanasundaram, Tugrul Bingol, Raghavan Muthuregunathan, Fedor Borisyuk

"arXiv:2608.28968v1 Announce Type: new Abstract: Semantic Search on LinkedIn must retrieve relevant profiles from a corpus of hundreds of millions in response to natural-language queries such as "a fintech founder in Berlin who worked in payments." The deployed relevance policy is…"

View on X

Originally posted by Dhritiman Das, Chujie Zheng, Ronak Kaoshik, Pratik Dixit, Vishal Shah, Yanbo Li, Jiahao Xu, Manika Agarwal, Chinmay Naik, Lingyu Zhang, Chetan Bhole, Chirag Bhanuprasad Mehta, Meng Zheng, Puneet Singh Ahluwalia, Shirisha Singh, Ping Jin, Manas Apte, Gokulraj Mohanasundaram, Tugrul Bingol, Raghavan Muthuregunathan, Fedor Borisyuk on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses