lixiaoxi45 commited on
Commit
1d5a947
·
verified ·
1 Parent(s): 45c2a2a

Upload folder using huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +2 -7
README.md CHANGED
@@ -22,8 +22,8 @@ license: apache-2.0
22
 
23
  WebThinker-R1-32B is part of the WebThinker series that enables large reasoning models to autonomously search, explore web pages, and draft research reports within their thinking process. This 32B parameter model provides deep research capabilities through:
24
 
25
- - **Deep Web Explorer**: Enables autonomous web searches and page navigation by clicking interactive elements to extract relevant information while maintaining reasoning coherence
26
- - **Autonomous Think-Search-and-Draft**: Integrates real-time knowledge seeking with content creation, allowing the model to draft sections as information is gathered
27
  - **RL-based Training**: Leverages iterative online DPO training with preference pairs constructed from reasoning trajectories to optimize end-to-end performance
28
 
29
  ## Related Models
@@ -40,11 +40,6 @@ This model can be used for:
40
  - Scientific research report generation
41
  - Open-ended reasoning tasks
42
 
43
- ## Resources
44
-
45
- - [GitHub Repository](https://github.com/RUC-NLPIR/WebThinker)
46
- - [Paper](https://arxiv.org/abs/2504.21776)
47
-
48
  ## Citation
49
 
50
  ```bibtex
 
22
 
23
  WebThinker-R1-32B is part of the WebThinker series that enables large reasoning models to autonomously search, explore web pages, and draft research reports within their thinking process. This 32B parameter model provides deep research capabilities through:
24
 
25
+ - **Deep Web Exploration**: Enables autonomous web searches and page navigation by clicking interactive elements to extract relevant information while maintaining reasoning coherence
26
+ - **Autonomous Think-Search-and-Draft**: Integrates real-time knowledge seeking with report generation, allowing the model to draft sections as information is gathered
27
  - **RL-based Training**: Leverages iterative online DPO training with preference pairs constructed from reasoning trajectories to optimize end-to-end performance
28
 
29
  ## Related Models
 
40
  - Scientific research report generation
41
  - Open-ended reasoning tasks
42
 
 
 
 
 
 
43
  ## Citation
44
 
45
  ```bibtex