News
· 2026-09-30
AWS has announced the release of Amazon Bedrock Knowledge Bases, a Retrieval Augmented Generation (RAG) technology that enables businesses to build conversational assistants capable of understanding natural language queries and providing cited answers. By parsing, chunking, embedding, and indexing documents, this technology offers an efficient solution for complex scenarios like claims processing. Amazon Bedrock Knowledge Bases supports multi-turn conversations, metadata filtering, and citation
News
· 2026-09-30
ArXiv has released new research on Retrieval-Augmented Generation (RAG) systems, introducing the Trajectory Variance Score (TVS) to detect semantic divergence caused by conflicts between retrieved context and parametric knowledge in RAG systems. TVS calculates the mean pairwise cosine distance of answer embeddings across independent stochastic denoising trajectories to capture the temporal tug of war between parametric and contextual attractors. Experiments across multiple datasets demonstrate t
News
· 2026-09-30
ArXiv has released a study on Retrieval-Augmented Generation (RAG) systems, focusing on the impact of document parsers, chunking strategies, and embedding models on RAG performance. The research conducted a controlled factorial study evaluating three parsers, three chunking strategies, and five dense embedding models, along with a sparse BM25 baseline, using 800 question instances across four distinct Indian government regulatory documents. The findings indicate that parser and chunker choices s
News
· 2026-09-29
The Hugging Face team introduces the Retrieval-Augmented Skill Optimization (RASO) framework, which leverages an external skill corpus as prior knowledge to optimize agent skills. RASO adapts relevant knowledge from existing skills to the target task and harness through Cross-Harness Adaptation, addressing mismatches in both domain and constraints. It consists of two stages: Retrieval-Augmented Skill Initialization (RASI) and Retrieval-Augmented Skill Update (RASU), enabling efficient skill cons
· 2026-09-29
This article explores the process of fine-tuning the Qwen3.5-4B model to optimize a Retrieval-Augmented Generation (RAG) pipeline, assessing its potential to replace Gemini Flash-Lite. Qwen3.5-4B, an open-source large language model, offers advantages in local deployment and resource optimization. By fine-tuning the model, researchers aim to enhance the efficiency and accuracy of the RAG pipeline in handling complex tasks, providing developers with a cost-effective and efficient AI solution.
News
· 2026-09-29
A new research paper on arXiv introduces the Trajectory Variance Score (TVS), a method for detecting knowledge conflicts in Retrieval-Augmented Generation (RAG) models within diffusion language models. TVS measures the temporal semantic divergence by computing the mean pairwise cosine distance of answer embeddings across independent stochastic denoising trajectories. Experiments across diverse datasets demonstrate that TVS can effectively identify and quantify conflicts between parametric knowle
News
· 2026-09-29
A new study published on arXiv presents a comprehensive evaluation framework for Retrieval-Augmented Generation (RAG) components, including 3 parsers, 3 chunking strategies, and 5 dense embedding models, along with a sparse BM25 baseline. The research, conducted across 800 question instances from four distinct Indian government regulatory documents, reveals that no single retriever family dominates across all documents, and the choice of parser and chunker significantly impacts performance. The
· 2026-09-25
Apowerb is an open-source AI agent runtime platform that focuses on supporting key features such as Retrieval-Augmented Generation (RAG), Text-to-SQL, and Webhooks. The platform aims to provide developers with flexible and scalable tools for building AI agents, streamlining the automation of complex tasks. The release of Apowerb opens new possibilities for AI agents in various application domains, particularly in data processing, automated querying, and system integration.
· 2026-09-22
AI·rete·RAG is a novel decision support system designed to address the issue of unaccountable AI decisions in critical areas such as lending, fraud detection, and clinical triage. It combines a pure-Python Rete rule engine with Retrieval-Augmented Generation (RAG) technology. The Rete engine evaluates YAML rules against facts to make decisions, while RAG retrieves relevant passages from policy documents and generates plain-English explanations for the decisions. This approach ensures traceabilit
News
· 2026-09-22
AdaMem is a relevance-guided soft compression framework designed to optimize memory allocation strategies in Retrieval-Augmented Generation (RAG) systems. By mapping learned passage-relevance estimates to a fixed memory-token budget, AdaMem enhances generation quality while reducing inference latency. Across six open-domain QA benchmarks, AdaMem achieved up to 3.2% substring match improvement under standard 16× compression and an average relative gain of 14.6% under aggressive 64× compression. A
News
· 2026-09-16
NepKANUN is an AI-powered legal assistant tailored for Nepali legal texts, built on a Retrieval-Augmented Generation (RAG) framework. It addresses the challenges of accessing legal information in Nepal, such as complex terminology, limited resources, and misinformation. Trained using a custom dataset of high-quality question-answer pairs, NepKANUN achieves strong F1 scores of 0.82 (simple), 0.77 (moderate), and 0.71 (complex) according to BERTScore. Expert reviews confirm its usability. This tec
News
· 2026-09-16
This study introduces a crash narrative-guided retrieval-augmented generation (RAG) framework that translates unstructured crash narratives into specific safety countermeasure recommendations. By extracting key attributes such as traffic control, signal indications, driver faults, and vehicle movements from crash descriptions, the framework leverages evidence-based treatments from the FHWA to provide actionable recommendations. Evaluated on 312 fatal and serious crashes across 115 intersections
· 2026-09-15
Make0 AI has launched Brain, a cross-agent shared memory system designed to address the issue of repetitive context initialization in AI agents during multitask processing. Utilizing RAG retrieval techniques and graph-based storage, Brain enables efficient and flexible memory management, supporting long-term memory sharing and interaction between AI agents. The system simplifies user operations with a drag-and-drop interface and visual tools, enhancing user experience. Its performance is impress
News
· 2026-09-14
LlamaIndex has released its latest newsletter, highlighting the introduction of LlamaParse, a highly accurate OCR agent platform for parsing and extracting content from complex documents efficiently. Additionally, the newsletter covers several product updates, including revamped documentation, a new contribution board, testing of the Zephyr-7b-beta model, enhanced image captioning, prompt compression techniques, and new integrations with HuggingFace and Gradient AI. These updates aim to improve
News
· 2026-09-14
Anthropic has launched Claude 100k, an AI model with a 100k token context window, which is approximately 75k words—three times that of GPT-4 (32k) and 25 times that of ChatGPT. This advancement allows for processing up to 300 pages of text in a single inference call, significantly enhancing capabilities in complex document analysis and financial report processing. Our tests on SEC 10-K filings demonstrate Claude 100k's impressive holistic understanding, latency, and cost implications for long-te
News
· 2026-09-14
In the TWIML AI podcast, LlamaIndex founder Jerry Liu delves into how LlamaIndex is revolutionizing AI application development by connecting large language models (LLMs) to private data sources. The discussion covers LlamaIndex's core features, such as advanced retrieval mechanisms, multi-data source information synthesis, and automated query interfaces. Additionally, it explores LlamaIndex's latest advancements in agent interaction, automated data processing, and optimizing AI application devel
News
· 2026-09-13
LlamaIndex has released a major update introducing new features such as Knowledge Graph integration, OpenAI agent support, and the FLARE technique for long-form generation. The update also enhances the user experience for LLM applications by supporting inline citations for improved transparency and traceability, and integrates with Microsoft Guidance for structured outputs. Additionally, LlamaIndex strengthens tool-calling capabilities, supports advanced query planning, multi-router functionalit
News
· 2026-09-13
LlamaIndex has launched an innovative system combining Text2SQL and RAG (Retrieval-Augmented Generation) to enhance product review analysis. The system decomposes user queries into database and interpretation queries, leveraging an in-memory SQLite database and the NLSQLTableQueryEngine for precise data retrieval. It then utilizes RAG to generate comprehensive answers. This approach streamlines complex query processing, improving AI application efficiency and accuracy in handling multi-dimension
News
· 2026-09-13
LlamaIndex has released a comprehensive guide on Retrieval Augmented Generation (RAG) systems, aiming to address the issue of outdated knowledge in large language models like ChatGPT. The guide explains that RAG enhances the accuracy and timeliness of generated responses by searching for relevant data and providing it to the LLM. LlamaIndex highlights RAG as an effective solution to the high costs and data limitations associated with updating LLM knowledge. The guide covers various technical app
News
· 2026-09-13
LlamaIndex has released a comprehensive guide on fine-tuning embedding models using synthetic data to enhance the performance of RAG (Retrieval-Augmented Generation) systems. This approach, which eliminates the need for manual labeling, leverages an LLM to generate hypothetical questions and relevant text pairs, resulting in a 5%-10% improvement in retrieval performance metrics. By optimizing the embedding space, the method enables models to better understand the semantics of specific domains, t
News
· 2026-09-13
In its October 2023 update, LlamaIndex introduced Multi-Document Agents (V1), a significant upgrade to its RAG (Retrieval-Augmented Generation) system. This new feature enables intelligent retrieval and asynchronous query planning across multiple documents, significantly enhancing the efficiency of complex document processing and the analytical capabilities of AI applications. Additionally, LlamaIndex has optimized integrations with platforms like HuggingFace and introduced new tools such as Uns
News
· 2026-09-13
LlamaIndex has released its September 2023 update, featuring a range of AI tool and framework enhancements, including Graph RAG, Neo4j integration, and the Sweep AI code splitter. These updates aim to improve the efficiency and accuracy of RAG (Retrieval-Augmented Generation) systems and support intelligent processing of complex and multimodal data. Additionally, LlamaIndex has introduced collaborative features with partners like Mendable AI and Nomic AI, providing developers with more powerful
News
· 2026-09-13
LlamaIndex has introduced the EmbeddingAdapterFinetuneEngine, a novel tool for fine-tuning a linear adapter on top of query embeddings from any embedding model. This approach optimizes RAG (Retrieval-Augmented Generation) system performance by transforming queries without the need to re-embed documents. The adapter is compatible with various embedding models, including SBERT, OpenAI, and Cohere, and allows for retraining on changing data distributions without re-embedding existing documents.
News
· 2026-09-13
LlamaIndex has released its September 2023 update, featuring the open-sourcing of its RAG (Retrieval-Augmented Generation) framework SECInsights.ai, linear adapter fine-tuning capabilities, and integrations with platforms like Replit. The update also introduces Hierarchical Agents and Hybrid Search to enhance complex data processing and AI application development. Additionally, LlamaIndex provides comprehensive fine-tuning guides and RAG building tutorials to help developers leverage LLM technol
News
· 2026-09-13
Timescale has launched the Timescale Vector integration for LlamaIndex, enhancing PostgreSQL as a vector database for AI applications. This integration offers up to 3x faster similarity search, efficient time-based filtering, and improved Retrieval Augmented Generation (RAG) capabilities. By consolidating vector embeddings, relational data, and time-series data into a single PostgreSQL database, Timescale Vector simplifies AI application development and enhances system performance, providing dev
News
· 2026-09-13
LlamaIndex's October 2023 update introduces a suite of enhancements and new features for its RAG (Retrieval-Augmented Generation) system, including Multi-Document Agents, the RetrieverEvaluator module, and native support for HuggingFace Embeddings. These updates aim to improve the efficiency of complex document parsing, retrieval evaluation, and AI application development. Additionally, LlamaIndex has integrated Arize AI Phoenix for comprehensive observability and introduced the LongContextReord
News
· 2026-09-13
LlamaIndex has released a comprehensive guide on optimizing chunk sizes for RAG (Retrieval-Augmented Generation) systems. The guide focuses on evaluating the impact of different chunk sizes on system accuracy and response times, using metrics like Faithfulness and Relevancy. It provides practical steps and tools based on LlamaIndex to help developers fine-tune their RAG systems for optimal performance.
News
· 2026-09-13
LlamaIndex's November 2023 newsletter introduces several updates to its AI toolchain and frameworks, including LlamaIndex Chat, Evaluator Fine-Tuning, and ParamTuner. These tools aim to enhance the efficiency and accuracy of RAG (Retrieval-Augmented Generation) systems. Additionally, LlamaIndex has integrated the latest embedding models from CohereAI Embed v3 and Voyage AI and improved its multimodal data processing capabilities. These updates highlight LlamaIndex's ongoing innovation in AI tool
News
· 2026-09-13
LlamaIndex has launched FinSight, a financial analysis application that leverages RAG (Retrieval-Augmented Generation) technology to streamline the analysis of company annual reports. Built on the Streamlit framework and powered by OpenAI's GPT-4, FinSight enables users to extract key financial insights and generate comprehensive reports efficiently. This release highlights the growing potential of AI in financial data analysis, providing investors and analysts with a powerful tool for informed
News
· 2026-09-13
LlamaIndex's October 2023 newsletter introduces a series of updates to its AI toolchain, including new features like QueryFusionRetriever, Router Fine-Tuning, and SQLRetriever, aimed at enhancing the efficiency and accuracy of RAG (Retrieval-Augmented Generation) systems. Additionally, LlamaIndex has expanded its LLM compatibility to include Amazon Bedrock and AI21 Labs and launched a multimodal RAG framework supporting intelligent retrieval and generation of both text and images. These updates