Repository navigation
原控制QA数量参数已不存在 || The original control QA quantity parameter no longer exists #187
Description
Activity
- changed the title
[-]原控制QA数量参数已不存在[/-][+]原控制QA数量参数已不存在 || The original control QA quantity parameter no longer exists[/+]on Apr 13, 2026 现在的参数配置示例见 https://github.com/InternScience/GraphGen/tree/main/examples/generate。
控制QA对数量取决于你使用哪种 partitioner——partition是指将 KG 转为很多子图的过程。
以 https://github.com/InternScience/GraphGen/blob/main/examples/generate/generate_aggregated_qa/aggregated_config.yaml 为例:- id: partition op_name: partition type: aggregate dependencies: - judge params: method: ece # ece is a custom partition method based on comprehension loss method_params: max_units_per_community: 20 # max nodes and edges per community min_units_per_community: 5 # min nodes and edges per community max_tokens_per_community: 10240 # max tokens per community unit_sampling: max_loss # unit sampling strategy, support: random, max_loss, min_loss可以通过设置:
- max_units_per_community:一个子图中最多有的节点或者边的数量
- min_units_per_community:一个子图中最少有的节点或者边的数量
- max_tokens_per_community:子图中节点或边的描述token数量总和最大值
来控制最终合成的QA数量(一般一个子图对应一个QA)。
我会尽快更新文档来说明最新版本的参数使用方法。
Reacted by HuHaibo现在的参数配置示例见 https://github.com/InternScience/GraphGen/tree/main/examples/generate。 控制QA对数量取决于你使用哪种 partitioner——partition是指将 KG 转为很多子图的过程。 以 https://github.com/InternScience/GraphGen/blob/main/examples/generate/generate_aggregated_qa/aggregated_config.yaml 为例:
- id: partition op_name: partition type: aggregate dependencies: - judge params: method: ece # ece is a custom partition method based on comprehension loss method_params: max_units_per_community: 20 # max nodes and edges per community min_units_per_community: 5 # min nodes and edges per community max_tokens_per_community: 10240 # max tokens per community unit_sampling: max_loss # unit sampling strategy, support: random, max_loss, min_loss可以通过设置:
- max_units_per_community:一个子图中最多有的节点或者边的数量
- min_units_per_community:一个子图中最少有的节点或者边的数量
- max_tokens_per_community:子图中节点或边的描述token数量总和最大值
来控制最终合成的QA数量(一般一个子图对应一个QA)。
我会尽快更新文档来说明最新版本的参数使用方法。
Thanks for your prompt reply.
如果我的理解是正确的话:每个文件产生的entities,即知识图谱的大小是取决于文件语义丰富度的。所以根据以上参数,能够根据文档大小自适应QA数量?即能保证QA的质量?
@AlbertLin0
知识图谱的大小是取决于文本语义丰富度。
以上参数可以调整 partition 这个过程(将整个KG划分为n个subgraph)中,每个subgraph蕴含的entities和relations的数量,由于每个subgraph 后面对应一个 QA,所以可以间接调整 QA 数量。
GraphGen 暂时没有提供直接调整 QA 数量的参数。如果需要的话,可以先自行在生成的 QA 中进行采样。
@AlbertLin0
The size of the knowledge graph depends on the semantic richness of the text.
The above parameters can adjust the number of entities and relations contained in each subgraph in the partition process (dividing the entire KG into n subgraphs). Since each subgraph corresponds to a QA, the number of QA can be adjusted indirectly.
GraphGen currently does not provide parameters for directly adjusting the number of QA. If desired, you can sample the generated QA yourself first.Reacted by HuHaibo
#28 以及 #27 控制QA对数量的方法,在代码中已经弃用了吗?
在config file中,已经找不到以下参数:
我在57页的文档默认生成了3w+个QA pairs。现版本有显式的参数控制生成数量吗?
Are #28 and #27 methods of controlling the number of QA pairs deprecated in the code?
In the config file, the following parameters can no longer be found:
My 57-page document generated 30,000+ QA pairs by default. Does the current version have explicit parameters to control the number of generations?