Nicolay Gerold

Nicolay Gerold

@nicolaygerold · Twitter ·

"Sadly, it's a bit off a snake oil. These long context embedding models have tested basically all of them, not really working well. So it's [best length of chunks] something between like 500 and 1,000 tokens." Text embeddings are far from perfect. They struggle with long documents. They struggle with out-of-domain-data. They struggle to incorporate additional factors like recency or trustworthiness. How can we fix this? - Fine-tuning helps adapt embeddings to specific domains, but demands careful data selection. Re-ranking aids in managing long documents and adds factors like recency and trustworthiness. - Re-ranking helps in managing long documents and can consider factors like recency and trustworthiness. In today’s episode on How AI Is Built, we are talking to @Nils_Reimers. He is one of the researchers who kickstarted the field of dense embeddings, developed sentence transformers, started @huggingface`s Neural Search team and now leads the development of search foundational models at @cohere. Tbh, he has too many accolades to count off here. Some insights from the episode: 1️⃣ Evolution of Embedding Models: - Shift from classification to contrastive learning - Importance of in-batch negatives and hard negatives - Scale batch size for better performance - Use of encoder-only models for text embeddings 2️⃣ Adapting Embeddings: - Use GenQ (Generate Queries) approach - Fine-tune models on domain-specific data - Continuous fine-tuning for evolving data 3️⃣ Enhancing Search: - Implement two-stage retrieval with re-ranking - Incorporate multiple factors (relevance, recency, trustworthiness) Link to the full episode is in the tread. ♻️ Repost this if you found something useful. #rag #embeddings #llm

Post media