Overview
These are the main pipeline settings. Experiment to find the best values for your use case.MATTER_CONFIG
The core pipeline object which determines which matter pages to process.
What are matter pages?
Default configuration
Why are title and author not in the fields section?Title and at least one author are always required and are implicitly included.
In other words, we need some baseline fields to name a file.This will be made configurable in the next update.
Front Matter
object
required
Front Matter Examples
Back matter
Back matter generally has less metadata we need and is designed primarily as a fallback if the title and at least one author name was not found in the front matter. There are certain situations where you may want to process back matter differently.object
required
Back Matter Examples
Counting Pages
number
default:"1"
required
Determines how to express page numbering when naming diagnostic files. 1-based makes
it easier to cross-reference page numbers to the source PDF page numbers.
Writing Changes
Critical setting that controls whether generated metadata and filenames are written to PDFs in[project-root]/data/ which are not in a
RUNTIME_IGNORE_DIR_NAME.
Only set this to True once you have run the script and confirmed suggestions are
acceptable. Review Evaluation for a full guide.
boolean
default:"false"
required
--ann fits for both the markdown and JSON exports.
number
default:"250"
required
Data folder
Drop your PDFs in the[project-root]/data folder. Read about
ignoring directories, a helpful feature for
troubleshooting.
The
[project-root]/data folder location is currently hardcoded. While this makes it
easy to ignore with git, it could be made configurable.This is planned for a future update. For now, you can create a symlink from your preferred
location to [project-root]/data.