There are many ways to integrate an industry’s knowledge into an LLM, but weighing cost and benefit, most enterprise applications use RAG. Recently, combining knowledge graphs (KGs) with RAG has become more and more common, and looks set to be the next hot technology. Riding that wave, on July 2 Microsoft finally open-sourced its GraphRAG, nearly five months after publishing the technical paper. I was very excited to test how well it works, so this article introduces GraphRAG’s concept, advantages and implementation steps, along with several ways to visualize it, in the hope of helping you put the technology into practice quickly.
1. What is GraphRAG?
The core idea of GraphRAG is to combine data stored in a knowledge graph with an LLM, improving the quality of the model’s answers by querying a graph database.
A knowledge graph is a structured way of representing knowledge, made up of entities and the relationships between them, so you can see how entities are connected on the graph, as in these conceptual examples:
If you want to go deeper into graphs and KGs, I recommend checking out Neo4j’s GraphAcademy or taking the Knowledge Graphs for RAG course on DeepLearning.AI. They cover RAG and KGs and will get you up to speed on the field quickly.
2. Why use GraphRAG?
Knowledge graphs go back at least to 2012, when Google launched its second-generation search engine and published a famous blog post, “Introducing the Knowledge Graph: things, not strings.” They found that organizing the information on web pages with a KG made a huge leap in capability possible. The same pattern is now developing in GenAI, where many projects hit a wall because they’re limited to processing strings rather than things. Today, frontier AI engineers and academic researchers are discovering the same important secret Google once found: the key to breaking through RAG’s bottlenecks lies in KGs, bringing knowledge of concrete things into statistics-based text techniques.
According to Microsoft’s GraphRAG research (Darren et al., 2024), GraphRAG greatly improves the retrieval part of traditional RAG, filling the context window with more relevant content, giving better answers and capturing the sources of evidence. They also found that GraphRAG needs 26% to 97% fewer tokens than other approaches, making it more scalable.
Here’s an example from the Microsoft Research Blog, where you can see GraphRAG’s answers are clearly better than ordinary RAG’s.
Another notable example comes from Writer. They recently published a RAG benchmark report based on the RobustQA framework (Mozolevskyi & AlShikh, 2024), comparing their GraphRAG approach with top competitors’ tools. As the results below show, GraphRAG scored 86%, far above the competitors’ range (33% to 76%), with comparable or lower latency.
In short, GraphRAG has three main advantages over traditional RAG:
- Higher accuracy and more complete answers, which helps a lot at runtime and in production.
- Once the KG is built, RAG applications are easier to build and maintain, a significant advantage in development time.
- Better explainability, traceability and access control, which benefits governance.
3. How do you use GraphRAG?
There are now many frameworks for implementing knowledge graph + RAG, such as LlamaIndex’s Property Graph Index, LangChain’s integration with Neo4j, and Haystack. The field is moving fast and the methods are getting easier to use. Today I’m introducing Microsoft GraphRAG, open-sourced not long ago.
GraphRAG is open source on GitHub, and they’ve written detailed documentation if you’re interested.
And if you don’t want to spend money but want to see what GraphRAG can do, I’ve put already-indexed files on GitHub: GraphRAG-Visualization-Tutorial, which you can download and use directly.
Step 1. Environment setup
It requires Python 3.10 to 3.12. I created the environment with conda.
conda create -n GraphRAG python=3.10
conda activate GraphRAG
Next, install GraphRAG
pip install graphrag
Step 2. Prepare the folder and documents
First, create a directory for documents, /ragtest/input. Then download the official reference document. You can download it here, or use the commands below to create the folder and download it.
Note that Microsoft GraphRAG currently supports only .txt; .pdf files are ignored.
mkdir -p ./ragtest/input
curl https://www.gutenberg.org/cache/epub/24022/pg24022.txt > ./ragtest/input/book.txt
Step 3. Initialize the workspace
Now we need to do some initial configuration. Here I’m using the default configuration; for details, see the official documentation.
python -m graphrag.index --init --root ./ragtest
Once it finishes, you’ll see that besides the input folder, the ragtest folder now has some new things in it.
Next, edit 2 files
.env: enter your OpenAI API or Azure OpenAI API key
GRAPHRAG_API_KEY=<API_KEY>
settings.yaml: lets you customize the whole pipeline, including which LLM to use. I’m using the defaults: the default embedding model is text-embedding-3-small, and the LLM is gpt-4-turbo-preview.
Step 4. Run the indexing pipeline
With the environment set up, and the .txt files you want indexed in the input folder, run the command below.
Note: this stage calls OpenAI’s API and costs a fair amount. If you want to skip it, you can download my repo, which has already been run.
python -m graphrag.index --root ./ragtest
It takes about 14 minutes. When it’s done, you’ll see the message Completed successfully.
Step 5. Ask questions
Now we can ask GraphRAG questions. GraphRAG’s query engine has three parts: local search, global search and question generation; for details, see the official documentation.
Here’s an example of asking an advanced question with global search:
python -m graphrag.query \
--root ./ragtest \
--method global \
"What are the top themes in this story?"
The answer took about 50 seconds. Here’s GraphRAG’s answer:
4. Visualization
As mentioned earlier, a KG is a graph showing the relationships between entities. There are a few ways to visualize the GraphRAG we built above.
Method 1. Use yFiles Graphs
We can use yFiles Graphs to help visualize GraphRAG, mainly by running a Jupyter notebook. I’ve uploaded the notebook, written by Microsoft, to GitHub. Before using it, install the packages in requirements.txt:
pip install requirements txt
After opening graph-visualization.ipynb, remember to change the directory paths below. The part after output is the time of your indexing run. You can also download my repo and use the defaults.
INPUT_DIR = "output/20240719-162300/artifacts"
Then follow the instructions and you’ll see a visualization of your KG:
Method 2. Generate GraphML and use third-party software
Before running the indexing pipeline, you can edit settings.yaml as follows:
snapshots:
graphml: True
That way, an extra GraphML file is generated after indexing. GraphML is an open standard supported by many open-source tools, so you can open it with software like Gephi or yEd Graph Editor.
5. How much does GraphRAG cost?
The document used above was downloaded from the official site and is about 185 KB. In OpenAI’s dashboard you can see this run cost about $5.25, which is actually pretty high for a file that size.
If you want to save money, I recommend changing the LLM in settings.yaml. The default is gpt-4-turbo-preview; using gpt-4o saves about half. See the price table below.
5. Conclusion
In this article we covered GraphRAG’s concept, advantages, implementation steps and visualization methods, so you can understand and apply the technology more intuitively.
To sum up, combining KGs with RAG has become a key technology for modern enterprises to improve data processing and question-answering systems. Microsoft’s GraphRAG is at the forefront of the field, using structured knowledge graphs to significantly improve the accuracy and relevance of model answers. And as GenAI develops, this matters in applications where answer quality is critical, where explainability is needed for internal, external or regulatory stakeholders, and where privacy and security controls on data access are required.
So I’ll boldly predict that RAG combined with KGs will play an increasingly important role in data retrieval, question answering and intelligent applications, becoming an important force driving AI forward.
References
Edge, D., Trinh, H., Cheng, N., Bradley, J., Chao, A., Mody, A., Truitt, S., & Larson, J. (2024). From Local to Global: A Graph RAG Approach to Query-Focused Summarization. ArXiv, abs/2404.16130.
Xu, Z., Cruz, M.J., Guevara, M., Wang, T., Deshpande, M., Wang, X., & Li, Z. (2024). Retrieval-Augmented Generation with Knowledge Graphs for Customer Service Question Answering. ArXiv, abs/2404.17723.
Support
If this article helped you, or you’d like to encourage me to keep writing, you can clap for it. Thank you for your support!