xushuwen23 commited on
Commit
de4cf3a
·
verified ·
1 Parent(s): 7b413a7

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +103 -2
README.md CHANGED
@@ -4,6 +4,107 @@ base_model:
4
  - Qwen/Qwen2.5-7B-Instruct
5
  pipeline_tag: question-answering
6
  tags:
7
- - agent
 
 
 
8
  ---
9
- GraphWalker-7B.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
  - Qwen/Qwen2.5-7B-Instruct
5
  pipeline_tag: question-answering
6
  tags:
7
+ - LLM Agent
8
+ - Knowledge Graph
9
+ - Question Answering
10
+ - Reasoning
11
  ---
12
+
13
+ # GraphWalker-7B
14
+
15
+ [**📄 Paper (arXiv:2603.28533)**](https://arxiv.org/abs/2603.28533) | [**💻 GitHub**](https://github.com/XuShuwenn/GraphWalker) | [**🤗 Model**](https://huggingface.co/xushuwen23/GraphWalker-7B)
16
+
17
+ **GraphWalker-7B** is a specialized large language model fine-tuned from [Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) for **Agentic Knowledge Graph Question Answering (KGQA)**. GraphWalker treats multi-hop KGQA as a sequential graph-walking process, learning to navigate knowledge graphs via synthetic trajectory curriculum — achieving strong generalization with a single, compact 7B model.
18
+
19
+ ---
20
+
21
+ ## 🌟 Overview
22
+
23
+ Multi-hop KGQA requires reasoning over complex, interconnected entities across a knowledge graph. Existing approaches either rely on rigid retrieval pipelines or expensive multi-module LLM orchestration. **GraphWalker** addresses this with two core ideas:
24
+
25
+ 1. **Agentic Graph Walking:** The model acts as an agent that iteratively traverses a KG, selecting which edges to follow at each step based on the question and accumulated context — effectively decomposing multi-hop questions into a series of local, grounded decisions.
26
+
27
+ 2. **Synthetic Trajectory Curriculum (STC):** Instead of relying on expensive human-annotated reasoning chains, GraphWalker is trained on *synthetically generated* graph-walking trajectories. The curriculum is structured to progressively increase trajectory complexity, enabling the model to internalize robust multi-hop strategies.
28
+
29
+ Together, these designs allow GraphWalker-7B to outperform much larger models and complex multi-agent systems on standard KGQA benchmarks, while remaining efficient at inference time.
30
+
31
+ ---
32
+
33
+ ## 🔑 Key Features
34
+
35
+ - **Agentic KG Navigation:** Frames KGQA as an iterative, step-by-step graph traversal rather than a single-shot retrieval-and-generate task.
36
+ - **Synthetic Trajectory Curriculum:** Trains on automatically constructed walking trajectories with progressively increasing difficulty, eliminating the need for costly human annotation.
37
+ - **Lazy Greedy Search with Frontier Expansion:** An efficient beam-search-style graph traversal algorithm that maximizes information gain while keeping context size tractable.
38
+ - **Single-Model Efficiency:** Achieves competitive performance with a single 7B model, without multi-agent overhead.
39
+ - **Broad KG Compatibility:** Designed to generalize across standard KGQA benchmarks (e.g., WebQSP, CWQ, GrailQA).
40
+
41
+ ---
42
+
43
+ ## 🛠️ Usage
44
+
45
+ ### 1. Environment Setup
46
+
47
+ ```bash
48
+ pip install vllm transformers
49
+ ```
50
+
51
+ ### 2. Download the Model
52
+
53
+ ```bash
54
+ # Via huggingface-cli
55
+ huggingface-cli download <your-org>/GraphWalker-7B --local-dir ./GraphWalker-7B
56
+ ```
57
+
58
+ ### 3. Inference with vLLM (Recommended)
59
+
60
+ **Start the vLLM server:**
61
+
62
+ ```bash
63
+ vllm serve "./GraphWalker-7B" \
64
+ --host 0.0.0.0 --port 22240 \
65
+ --served-model-name graphwalker-7b \
66
+ --gpu-memory-utilization 0.9 \
67
+ --dtype auto \
68
+ --chat-template "./GraphWalker-7B/chat_template.jinja"
69
+ ```
70
+
71
+
72
+
73
+ ---
74
+
75
+ ## 📈 Evaluation Results
76
+
77
+ | Method | Backbone | CWQ EM | CWQ F1 | WebQSP EM | WebQSP F1 |
78
+ |:---|:---|:---:|:---:|:---:|:---:|
79
+ | **GraphWalker** | | | | | |
80
+ | †Vanilla Agent | Qwen2.5-7B-Instruct | 40.7 | 33.2 | 68.4 | 66.1 |
81
+ | †Vanilla Agent | GPT-4o-mini | 63.4 | 60.3 | 79.6 | 70.6 |
82
+ | †Vanilla Agent | DeepSeek-V3.2 | 69.8 | 63.5 | 76.7 | 71.8 |
83
+ | GraphWalker-7B-SFT | Qwen2.5-7B-Instruct | 68.3 | 63.2 | 82.0 | 79.1 |
84
+ | GraphWalker-3B-SFT-RL | Qwen2.5-3B-Instruct | 70.9 | 65.2 | 83.5 | 81.7 |
85
+ | GraphWalker-8B-SFT-RL | LLaMA3.1-8B-Instruct | <u>78.5</u> | 69.6 | <u>88.2</u> | <u>84.5</u> |
86
+ | **GraphWalker-7B-SFT-RL** | **Qwen2.5-7B-Instruct** | **79.6** | **74.2** | **91.5** | **88.6** |
87
+
88
+ ---
89
+
90
+ ## 📝 Citation
91
+
92
+ If you use GraphWalker-7B or find this work helpful, please cite:
93
+
94
+ ```bibtex
95
+ @misc{xu2026graphwalkeragenticknowledgegraph,
96
+ title={GraphWalker: Agentic Knowledge Graph Question Answering via Synthetic Trajectory Curriculum},
97
+ author={Shuwen Xu and Yao Xu and Jiaxiang Liu and Chenhao Yuan and Wenshuo Peng and Jun Zhao and Kang Liu},
98
+ year={2026},
99
+ eprint={2603.28533},
100
+ archivePrefix={arXiv},
101
+ primaryClass={cs.CL},
102
+ url={https://arxiv.org/abs/2603.28533},
103
+ }
104
+ ```
105
+
106
+ ---
107
+
108
+ ## 📄 License
109
+
110
+ This model is released under the [Apache 2.0 License](https://www.apache.org/licenses/LICENSE-2.0), consistent with the base model [Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct).