RepoDaily · 2026-08-14 · Infrastructure / Runtime

RAGFlow pairs template-based document chunking with agent orchestration across a dozen model providers

#11 Infrastructure / Runtime Go +473 infiniflow/ragflow Open repository

infiniflow's open-source RAG engine fuses deep document parsing, GraphRAG, and MCP-based agents under Apache-2.0, requiring Python 3.13 and Go 1.26.4.

Repo typeInfrastructure / Runtime
Best forEngineering teams building retrieval-augmented Q&A systems over complex document formats that need precise chunking, citation tracking, and multi-provider LLM agent orchestration
Risk levelMedium — large dependency surface, a documented pickle deserialization vulnerability, and a stale supported-versions table in SECURITY.md
Time to evaluate2 to 3 days to pull the Docker image, ingest a representative document corpus, and compare retrieval quality against a baseline

Primary question: Does RAGFlow's template-aware chunking quality justify operating its multi-service stack versus wiring together a lighter RAG library?

91/100

RepoDaily adoption score

RepoDaily rates this as 91/100 (strong) for adoption: evidence, installation path, production risk, differentiation, license clarity, and AI/agent fit are scored from the article sources and adoption notes.

Directional score from RepoDaily sources and adoption notes, not a benchmark.Risk: Medium
100Evidence quality

5 source(s) across 3 source category/categories, plus a RepoDaily-specific evidence module when available.

100Installability

5 workflow step(s), 5 next-action step(s), and 4 command/install signal(s) were detected.

63Maintenance confidence

Trending momentum is +473 stars, with maintenance/release/issue signals counted when present.

93Production readiness

Risk is marked medium, with 5 security note(s) and 4 explicit skip condition(s).

100Differentiation

3 opportunity lens item(s), 4 alternative(s), and 4 type-specific section(s) support differentiation.

82License clarity

License source or license wording is present.

84Agent / AI fit

6 AI/agent-related signal(s) were detected in the article text and metadata.

Project overview

RAGFlow (v0.26.4) is an open-source Retrieval-Augmented Generation engine developed by infiniflow. The project description positions it as combining deep document understanding with agent capabilities to produce a context layer for large language models. The repository is primarily categorized as Go in GitHub's metadata, but pyproject.toml reveals a substantial Python application layer with over 100 pinned dependencies spanning document parsing, embedding, vector storage, and LLM client libraries.

The project's pyproject.toml declares a strict Python version constraint of >=3.13,<3.14, while go.mod specifies Go 1.26.4. This dual-runtime architecture means the Go layer (built with Gin, GORM, Redis, and NATS per go.mod) likely handles API serving, authentication, and orchestration, while the Python layer handles document ingestion, embedding generation, and LLM interaction. The dependency list in pyproject.toml includes clients for Anthropic (0.76.0), Cohere (5.6.2), Mistral, Groq, DashScope (Alibaba/Qwen at 1.25.11), Qianfan (Baidu at 0.4.6), Ollama, and Google GenAI, giving RAGFlow out-of-the-box compatibility with most major model providers.

RAGFlow also integrates agentic capabilities through the MCP protocol (mcp>=1.28.1), browser-use (>=0.11.1), and crawl4ai (>=0.9.2), positioning it beyond static document retrieval into agent-driven research. The inclusion of graspologic (pinned to a specific Gitee commit) suggests GraphRAG functionality for knowledge-graph-augmented retrieval. The README highlights recognition in the GitHub Octoverse and Trendshift, and the project maintains documentation at ragflow.io/docs/dev/, a public roadmap at GitHub issue 12241, and a Discord community.

Security posture requires attention. SECURITY.md documents a code-execution vulnerability in the restricted_loads function at api/utils/__init__.py line 215, where the numpy.f2py.diagnose.run_command function bypasses the intended pickle deserialization allowlist. Additionally, the supported-versions table in SECURITY.md lists only versions <=0.7.0 as receiving security updates, which appears stale relative to the current 0.26.4 release.

Problem it solves

  • Naive text chunking splits tables, forms, and multi-column layouts, producing retrieval fragments that strip critical context
  • Hallucination in LLM answers when retrieval context is shallow or poorly grounded in source documents
  • Vendor lock-in from proprietary RAG platforms that limit model choice and prevent on-premise deployment
  • Fragmented tooling when teams need to combine document parsing, vector indexing, agent orchestration, and multi-provider LLM access
  • Difficulty tracking which source passage generated each claim, undermining trust in retrieval-augmented answers

How it works

  1. Documents are ingested through the Python pipeline, which uses pdfplumber (0.11.10), python-docx, python-pptx, and opencv-python (4.10.0.84) to parse complex formats including PDF, Office documents, and images
  2. Parsed content is segmented using template-aware chunking strategies rather than fixed-size splits, as demonstrated in the README's chunking visual asset
  3. Chunks are embedded via ONNX Runtime (1.23.2 for GPU or CPU depending on platform) or provider-specific embedding APIs, then indexed in Elasticsearch (elasticsearch-dsl 8.12.0), OpenSearch (2.7.1), or InfiniFlow's own Infinity vector database (infinity-sdk 0.7.3)
  4. Agent orchestration layers use MCP (>=1.28.1), browser-use, and crawl4ai to extend retrieval with live web access and tool-calling capabilities
  5. At query time, the Go API layer (Gin-based) routes requests through retrieval, ranking, and LLM generation, returning answers with citations back to specific document chunks

Product demo and interface preview

Chunking demonstration
Chunking demonstration — Shows how RAGFlow segments a complex document into retrievable chunks with visual bounding boxes — the project's core differentiator versus naive text splitting. README.md image
Agentic workflow demonstration
Agentic workflow demonstration — Illustrates the agent orchestration layer where tools, retrieval, and LLM reasoning are chained — reflecting the mcp and browser-use dependencies in pyproject.toml. README.md image
RAGFlow in the GitHub Octoverse
RAGFlow in the GitHub Octoverse — GitHub Octoverse recognition displayed in the project README, indicating visibility within the open-source RAG category. README.md image

Architecture: Go gateway over a Python processing core

go.mod declares module 'ragflow' with go 1.26.4 and pulls in Gin (v1.12.0) for HTTP routing, GORM (v1.25.7) with MySQL and SQLite drivers for relational storage, Redis (go-redis v9) for caching and session state, and NATS (v2.14.3 server with v1.52.0 client) for internal messaging. OpenTelemetry (v1.44.0) is wired in for tracing, and the AWS SDK v2 and Google Cloud Storage client are present for object storage.

pyproject.toml declares ragflow v0.26.4 requiring Python >=3.13,<3.14. Notable pinned dependencies include anthropic==0.76.0, cohere==5.6.2, dashscope==1.25.11, qianfan==0.4.6, ollama>=0.5.0, and google-genai>=1.41.0. The graspologic dependency is pinned to a specific commit on Gitee (infiniflow's mirror), indicating a custom fork for GraphRAG graph construction. The inclusion of both infinity-sdk (0.7.3) and infinity-emb (>=0.0.66) ties RAGFlow to InfiniFlow's own vector database and embedding server.

Integration surface: model providers, storage backends, and agent tools

  • Model providers: Anthropic, Cohere, Mistral, Groq, Alibaba DashScope (Qwen), Baidu Qianfan, Google GenAI, Ollama, and OpenAI-compatible APIs — all present as pinned or bounded dependencies in pyproject.toml
  • Vector stores: Elasticsearch (elasticsearch-dsl 8.12.0), OpenSearch (opensearch-py 2.7.1), and InfiniFlow Infinity (infinity-sdk 0.7.3)
  • Object storage: MinIO (7.2.4), AWS S3 (boto3), Google Cloud Storage, Azure Data Lake (12.16.0), Dropbox, and Box SDK
  • Agent tooling: MCP protocol client (mcp>=1.28.1), browser-use (>=0.11.1) for web automation, crawl4ai (>=0.9.2) for web crawling, and DuckDuckGo search (>=7.2.0)
  • Document formats: PDF (pdfplumber 0.11.10, pypdf), Word (python-docx), PowerPoint (python-pptx), Excel (python-calamine, openpyxl via excelize in Go), HTML (html-text, markdownify), and email (extract-msg)
  • Collaboration integrations: Discord (discord-py 2.3.2 with audioop-lts backport for Python 3.13), Jira (3.10.5), Atlassian API (4.0.7), DingTalk, and Lark/Feishu SDKs

Deployment notes: Docker-first with strict runtime requirements

The README links to a Docker image at infiniflow/ragflow on Docker Hub, currently tagged at v0.26.4. The Docker image badge in the README references docker pull infiniflow/ragflow:v0.26.4. Teams deploying from source must provision Python 3.13 exactly (the constraint is >=3.13,<3.14) and Go 1.26.4, which are both recent runtime versions as of mid-2026.

pyproject.toml notes a specific Python 3.13 compatibility issue: discord-py 2.3.2 imports the removed audioop module, requiring the audioop-lts (>=0.2.1) backport. The Werkzeug version is also constrained to avoid multipart parsing bugs. These pinned workarounds indicate the team actively maintains compatibility against bleeding-edge runtimes but teams adopting RAGFlow should expect to match these exact versions.

Maintenance risk: dependency breadth and security policy gaps

  • pyproject.toml lists over 100 direct dependencies — a surface area that increases transitive vulnerability exposure and upgrade coordination cost
  • SECURITY.md documents a code-execution vulnerability in restricted_loads (api/utils/__init__.py line 215) exploitable via numpy.f2py.diagnose.run_command, with a proof-of-concept provided
  • The supported-versions table in SECURITY.md only lists versions <=0.7.0 as supported, while the current release is 0.26.4 — this table appears outdated and creates ambiguity about which versions receive patches
  • graspologic is pinned to a specific Gitee commit (38e680cab72bc9fb68a7992c3bcc2d53b24e42fd) rather than a tagged release, which complicates reproducibility and upstream tracking

Who should pay attention?

Good fit if

  • Teams ingesting PDFs with tables, forms, or multi-column layouts where naive chunking produces unusable fragments
  • Organizations needing on-premise deployment with a choice of LLM providers including DeepSeek, Ollama, or Qwen
  • Projects that require GraphRAG or agent-driven research alongside static document retrieval
  • Engineering teams comfortable operating Docker, Elasticsearch or OpenSearch, and a Go-plus-Python codebase

Skip for now if

  • Simple FAQ chatbots that only need keyword search over plain text
  • Teams without containerization experience or dedicated infrastructure for Elasticsearch, Redis, and object storage
  • Projects that require a minimal dependency tree — RAGFlow pulls in over 100 Python packages alone
  • Environments locked to Python 3.11 or 3.12, since RAGFlow requires Python >=3.13,<3.14

Risks and cautions

Medium

The project is actively maintained at v0.26.4 with a public roadmap, but the dependency surface exceeds 100 Python packages, a documented pickle deserialization vulnerability exists, and the SECURITY.md supported-versions table is stale.

  • SECURITY.md documents a code-execution vulnerability in restricted_loads via numpy.f2py.diagnose.run_command with a working proof-of-concept
  • The supported-versions table in SECURITY.md lists only <=0.7.0 as supported, creating uncertainty about patch coverage for the current 0.26.4 release
  • Python runtime is pinned to a narrow >=3.13,<3.14 window, limiting compatibility with existing environments
  • graspologic is sourced from a specific Gitee commit rather than a PyPI release, complicating supply-chain auditing
  • The dual Go-and-Python architecture doubles the runtime and toolchain management burden compared to single-language alternatives
  • Licensed under Apache-2.0 per LICENSE, permitting commercial use, modification, and redistribution
  • SECURITY.md discloses a restricted_loads vulnerability at api/utils/__init__.py line 215 where numpy module imports bypass the pickle allowlist, enabling remote code execution
  • The proof-of-concept in SECURITY.md demonstrates executing arbitrary shell commands (whoami) through the numpy.f2py.diagnose.run_command function
  • SECURITY.md's supported-versions table only marks versions <=0.7.0 as receiving security updates, despite the current release being 0.26.4
  • Dependencies include pycryptodomex (3.20.0) and paramiko (>=3.5.1), which should be audited for transitive CVEs in production deployments

Alternatives to compare

ApproachWhen to useTrade-off
LlamaIndex
When you need a Python-native RAG framework with fine-grained control over indexing strategies and a large plugin ecosystemFree, MIT-licensed
Dify
When you want a visual pipeline builder for LLM applications with built-in RAG and agent orchestrationFree self-hosted, Apache-2.0; paid cloud tiers available
Haystack (deepset)
When you need production-grade search pipelines with strong typing and connectors to multiple vector storesFree, Apache-2.0
LangChain
When you need a general-purpose LLM application framework with broad community support and extensive integrationsFree, MIT-licensed

What this trend reveals

Enterprise document intelligence with precise chunking

RAGFlow's template-aware chunking, visible in the README's chunking demonstration, targets the gap where generic RAG libraries fail on structured documents like financial reports and technical specifications.

Run a side-by-side comparison of RAGFlow's chunking against fixed-size splitting on 50 complex PDFs and measure retrieval precision with a held-out question set.

Deep research agents with MCP integration

The mcp>=1.28.1 dependency and browser-use integration position RAGFlow for agent workflows that combine live web research with grounded document retrieval — a use case the topics list explicitly calls out as 'deep-research'.

Build a prototype agent that queries both the document index and live web sources via MCP, then evaluate answer quality against a research benchmark.

Multi-model private deployment

With Ollama, DashScope, Qianfan, and OpenAI-compatible endpoints all wired in, RAGFlow suits organizations that route different query types to different models for cost or latency optimization.

Configure two model backends (one local via Ollama, one cloud) and measure latency and cost per 1,000 queries on a representative workload.

Best next action

Containerize and benchmark chunking quality on your document corpus

Pull the official Docker image and test RAGFlow's template-based chunking against your most challenging document formats before committing to the full stack.

  1. Pull infiniflow/ragflow:v0.26.4 from Docker Hub and follow the self-hosting instructions in the README
  2. Ingest 20-30 representative documents (PDFs with tables, multi-column layouts, or mixed text-image pages)
  3. Review the chunking output in the RAGFlow UI, checking whether tables and structured sections remain intact
  4. Configure one LLM provider (Ollama for local or an API key for a cloud provider) and run 50 test queries
  5. Compare citation accuracy and answer quality against your current retrieval baseline

RepoDaily verdict

RAGFlow at v0.26.4 is a feature-dense RAG platform that distinguishes itself through template-aware chunking, GraphRAG via a custom graspologic fork, and agent orchestration through MCP. The trade-off is a heavy dependency surface (100+ Python packages, dual Go-and-Python runtime), a documented code-execution vulnerability in the pickle deserialization path, and a stale security-policy table. Teams that prioritize chunking quality and multi-provider flexibility over operational simplicity will find the most value here.

Sources