Skip to content

原控制QA数量参数已不存在 || The original control QA quantity parameter no longer exists #187

Description

@AlbertLin0

#28 以及 #27 控制QA对数量的方法,在代码中已经弃用了吗?

在config file中,已经找不到以下参数:

traverse_strategy:
  qa_form: aggregated
  bidirectional: true
  edge_sampling: max_loss
  expand_method: max_width
  isolated_node_strategy: ignore
  max_depth: 1
  max_extra_edges: 2
  max_tokens: 256
  loss_strategy: only_edge

我在57页的文档默认生成了3w+个QA pairs。现版本有显式的参数控制生成数量吗?


Are #28 and #27 methods of controlling the number of QA pairs deprecated in the code?

In the config file, the following parameters can no longer be found:

traverse_strategy:
  qa_form: aggregated
  bidirectional: true
  edge_sampling: max_loss
  expand_method: max_width
  isolated_node_strategy: ignore
  max_depth: 1
  max_extra_edges: 2
  max_tokens: 256
  loss_strategy: only_edge

My 57-page document generated 30,000+ QA pairs by default. Does the current version have explicit parameters to control the number of generations?

Activity

  1. changed the title [-]原控制QA数量参数已不存在[/-] [+]原控制QA数量参数已不存在 || The original control QA quantity parameter no longer exists[/+] on Apr 13, 2026
  2. ChenZiHong-Gavin commented on Apr 13, 2026

    @ChenZiHong-Gavin
    Collaborator

    现在的参数配置示例见 https://github.com/InternScience/GraphGen/tree/main/examples/generate。
    控制QA对数量取决于你使用哪种 partitioner——partition是指将 KG 转为很多子图的过程。
    以 https://github.com/InternScience/GraphGen/blob/main/examples/generate/generate_aggregated_qa/aggregated_config.yaml 为例:

      - id: partition
        op_name: partition
        type: aggregate
        dependencies:
          - judge
        params:
          method: ece # ece is a custom partition method based on comprehension loss
          method_params:
            max_units_per_community: 20 # max nodes and edges per community
            min_units_per_community: 5 # min nodes and edges per community
            max_tokens_per_community: 10240 # max tokens per community
            unit_sampling: max_loss # unit sampling strategy, support: random, max_loss, min_loss
    

    可以通过设置:

    • max_units_per_community:一个子图中最多有的节点或者边的数量
    • min_units_per_community:一个子图中最少有的节点或者边的数量
    • max_tokens_per_community:子图中节点或边的描述token数量总和最大值

    来控制最终合成的QA数量(一般一个子图对应一个QA)。

    我会尽快更新文档来说明最新版本的参数使用方法。

  3. ChenZiHong-Gavin commented on Apr 13, 2026

    @ChenZiHong-Gavin
    Collaborator
  4. AlbertLin0 commented on Apr 14, 2026

    @AlbertLin0
    Author

    现在的参数配置示例见 https://github.com/InternScience/GraphGen/tree/main/examples/generate。 控制QA对数量取决于你使用哪种 partitioner——partition是指将 KG 转为很多子图的过程。 以 https://github.com/InternScience/GraphGen/blob/main/examples/generate/generate_aggregated_qa/aggregated_config.yaml 为例:

      - id: partition
        op_name: partition
        type: aggregate
        dependencies:
          - judge
        params:
          method: ece # ece is a custom partition method based on comprehension loss
          method_params:
            max_units_per_community: 20 # max nodes and edges per community
            min_units_per_community: 5 # min nodes and edges per community
            max_tokens_per_community: 10240 # max tokens per community
            unit_sampling: max_loss # unit sampling strategy, support: random, max_loss, min_loss
    

    可以通过设置:

    • max_units_per_community:一个子图中最多有的节点或者边的数量
    • min_units_per_community:一个子图中最少有的节点或者边的数量
    • max_tokens_per_community:子图中节点或边的描述token数量总和最大值

    来控制最终合成的QA数量(一般一个子图对应一个QA)。

    我会尽快更新文档来说明最新版本的参数使用方法。

    Thanks for your prompt reply.

    如果我的理解是正确的话:每个文件产生的entities,即知识图谱的大小是取决于文件语义丰富度的。所以根据以上参数,能够根据文档大小自适应QA数量?即能保证QA的质量?

  5. ChenZiHong-Gavin commented on Apr 14, 2026

    @ChenZiHong-Gavin
    Collaborator

    @AlbertLin0
    知识图谱的大小是取决于文本语义丰富度。
    以上参数可以调整 partition 这个过程(将整个KG划分为n个subgraph)中,每个subgraph蕴含的entities和relations的数量,由于每个subgraph 后面对应一个 QA,所以可以间接调整 QA 数量。
    GraphGen 暂时没有提供直接调整 QA 数量的参数。如果需要的话,可以先自行在生成的 QA 中进行采样。


    @AlbertLin0
    The size of the knowledge graph depends on the semantic richness of the text.
    The above parameters can adjust the number of entities and relations contained in each subgraph in the partition process (dividing the entire KG into n subgraphs). Since each subgraph corresponds to a QA, the number of QA can be adjusted indirectly.
    GraphGen currently does not provide parameters for directly adjusting the number of QA. If desired, you can sample the generated QA yourself first.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions