Skip to content

Preprocess Office QA training dataset #151

Description

@tli2

TODO:

  • Understand types of questions asked (computation, visual understanding, etc.) -- try to cluster them into question templates
  • Extract shared computations that are common among multiple questions (i.e., index candidates)
  • Statistics about retrieval count (number of documents in Oracle, etc.)
  • Anything you might notice as you look through the dataset

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions