Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 

Repository files navigation

Gen's Magical Spreadsheets

A repo for the many beautiful spreadsheets created over the course of work with DDMAL! Most were created by Gen, to keep track of her various activities, but many other people have added and contributed to them. When adding more spreadsheets, please make sure to indicate who contributed and when (if known), as well as a description of what the spreadsheet contains (and each sheet, where applicable).

This is the full list of spreadsheets. Clicking on a spreadsheet name brings you to its description in this README, and from there to the spreadsheet itself.

Reference spreadsheets (these contain timeless information that could still be useful)

Neume notation examples

https://docs.google.com/spreadsheets/d/11Hfw4gD4au10WkWjswVRbE1t02NCWA3mBemxQ8zm7s0/edit?gid=0#gid=0

Images of the most common neume shapes in both Hufnagel (sheet 1) and square (sheet 2) notations. Can be useful to understand what each neume name means and how to recognize them!

CantusDB catalogue

https://docs.google.com/spreadsheets/d/17Nbc39AeAA12oF2Hk0G7wMaUYArXfsT_rEmnyoAcCeo/edit?gid=1855664281#gid=1855664281

A catalogue of all indexed sources on CantusDB (as of Summer 2026), made by Gen. Information for each manuscript includes CantusDB link, whether the source has images/text/volpiano on CantusDB, how many folios/pages/chants the source has, info on the script and/or notation style, provenance, and date.

Record-keeping spreadsheets - manuscript correction

MS073 - Jiali's magical spreadsheet 2.0

https://docs.google.com/spreadsheets/d/1wQcaxQh09q6TxPqFtbhse3KwAd-W6vZBmxo103B13-k/edit?gid=0#gid=0

This spreadsheet contains the record of the omr-ization and correction the MS73 manuscript in... 2024? It includes information about the different hands found in the manuscript, who corrected the folios, when, and how long it took. The spreadsheet was originally created by Jiali, but Gen (GGP) and Kyrie (KB) were the main contributors. Final mei files can be found here.

Another magical spreadsheet

https://docs.google.com/spreadsheets/d/1JjNyhZvEtv-3U5kgqcnsE0mCwHs8A-Px12CYYNLPsVw/edit?gid=0#gid=0

This spreadsheet contains the record of the omr-ization and correction of the Einsiedeln antiphoner (CH-E 611) in starting in 2022-23. It includes information about who corrected what, how long it took, and various problems we ran into. Gen, Ai Lynn, Zi Hang, Phoebe from Halifax, and Nicholas and Henry from Leuven worked on the initial correction, supervised by Gen. In 2024, when files were not all yet reviewed, we found a problem with C clef octaves, which led to us making various corrections to various files. A record of this is included in the spreadsheet as well. Final mei files can be found here.

Correcting stemmed notes in Salzinnes

https://docs.google.com/spreadsheets/d/1xfvc_T8lwRuT0DGmDc0rrJvj0545Bla9Z1ytVVZpZYw/edit?pli=1&gid=0#gid=0

When we first corrected Salzinnes, we made everything into puncta, because that's what we taught the IC to classify (thanks to the handy-dandy diagonal neume slicer). However, some time later, someone (I think Anna) pointed out that actually quite a few of these neumes had visible stems and were therefore virgas. We went back over all the corrected mei files to add stems where necessary and kept track of our work in this spreadsheet.

Gen's magical spreadsheet (the OG!)

https://docs.google.com/spreadsheets/d/1DmNAhP2MoIwyNP6EgA-YJr6xC03_V895tokWRCTMMu0/edit?gid=0#gid=0

This spreadsheet contains the record of the omr-ization and correction of the Salzinnes antiphoner (CDN-Hsmu M2149.L4 1554) in 2021 and 2022 (I think). It includes who corrected what and correction time. Gen, Yinan, Jiali and Andy worked on the initial correction, guided by Anna. Changes then happened within Rodan and Neon that made some elements of the Salzinnes mei files obsolete or incorrect; Gen came back in 2024 to finish reviewing files and make the necessary corrections. Final mei files can be found here.

Record-keeping spreadsheets - others

Mothra Bonanza (the big one)

https://docs.google.com/spreadsheets/d/1K_5yCKC0514RhpI1Lw0fEmRgGbwWakga1OaEQ3wt0qQ/edit?gid=0#gid=0

This spreadsheet contains records of various ground truth production and training runs performed in 2026 to get Mothra up and running. Everything was done by Gen unless otherwise specified.

First round of corrections: Folios with layers detected by the very first YOLO phase 1 detection model were sent to Gen for evaluation and correction. The corrected files were then used to improve the YOLO model.

Additional staff layers: At various points in time, more staff line ground truth was requested. Some was for a YOLO layer detection model, and some was for a Paco layer extraction model; this is made clear in the spreadsheet.

YOLO neume splitting GT: The initial YOLO phase 1 model didn't split neumes sufficiently for the IC to work efficiently. This ground truth was created in August 2026 to try to make a second YOLO model to split the neumes as we want. The sheet lists the folios used.

Additional neume layers: In August 2026, more neume detection ground truth was created to try to improve the YOLO phase 1 model. The sheet lists the folios used.

Pixel vs. Mothra showdown: This sheet describes tests performed early on to determine whether the new YOLO detection model could separate layers better than Rodan.

Method B annotations: This sheet was created by Cassie, I think.

Kraken line seg training: This sheet describes Gen's (fruitless) efforts to train Kraken to recognize and segment two-column manuscripts properly. Cassie eventually found a better solution.

CATDOES manuscripts: Made by Cassie to keep track of which manuscripts we're planning to run through CATDOES.

Output evaluations

https://docs.google.com/spreadsheets/d/1RxQUY4KNclOdGByYcqv1Iekw7na8bgdlSFulXluMcag/edit?gid=379490544#gid=379490544

A record of the various tests and evaluations Gen performed during Mothra creation and training times (summer 2026).

Paco staff detection: Determining which MS73 layer separation model, Betty or Eliza, worked best on various other manuscripts (answer: neither).

Baseline detection: Determining which line segmentation model (htrflow-rtmdet-lines, htrflow-yolov9, or kraken) worked best on the widest range of manuscripts (answer: kraken). It was suggested that this could be an interesting article, so info includes numbers and stats showing which model did best for each folio.

Mothra vs. standalone text comparison: For a little while, the text detection workflow in mothra was producing different results from the standalone text workflow. This sheet details Gen's attempts to figure out exactly where the differences lay, and which did better.

Mothra first round results: Describing the result of running Phase 1 YOLO layer detection on various folios for the first time. This was before the lab decided to use Paco to detect staff lines.

Masking tests: Made by Cassie to document which folios were used to test mothra-text output across corrected, uncorrected, and no-masking conditions.

UMIL Master Spreadsheet

https://docs.google.com/spreadsheets/d/1KTTa9WqfmU8narIXexLqXTYSl8MtkZfe/edit?pli=1&gid=42692703#gid=42692703

2025 was the UMIL year. Gen, with the help of Caroline and Pouya, looked for new translations of instruments on UMIL, as well as new instruments to add to UMIL, and kept track of their efforts here.

Welcome page: Explains what the spreadsheet is about, details contribution guidelines

Instruments we want to add: The very long list of instruments Gen and Caroline found that could be added to UMIL. Some were eventually added, most were not (this is made clear with the colours).

Rejected entries: A list of all the instruments that we came across and decided not to add to UMIL, and why.

Wikidata preselection list: Instruments already on Wikidata, but not on UMIL, that we wanted to fast-track to UMIL.

Problematic existing entries: Instruments already on UMIL that we wanted to remove or correct (reasons included).

Electrophones: Where Caroline put electrophones she didn't know what to do with yet.

Instruments no French: All the instruments on UMIL that didn't have a French label as of 17.07.25. Gen made the initial list and then Caroline spent more time on it and made it much better; their various discussions are included in the comments. Most labels have been added.

Instruments no German: All the instruments on UMIL that didn't have a German label as of 12.01.26. Some labels have been added.

Instruments no English: All the instruments on UMIL that didn't have an English label as of 23.07.25. Finding translations for these turned out to be very difficult, so most labels were not added.

Missed Iranian instruments: A list of Iranian instruments not yet on UMIL, made by Pouya.

Global Jukebox instruments: As described in the sheet: Instruments Gen found on Global Jukebox at the beginning of her search when she was doing a less good job.

Keeping track of Bach shenanigans

https://docs.google.com/spreadsheets/d/1Alns0hyCL_JCoCRI_J1aOxck8rw2YudANK5ejoF_Pgg/edit?gid=0#gid=0

In 2023-2024(?) Gen (GGP) and Kyrie (KEB) tried really really hard to teach Rodan to read Bach manuscript pages and extract the clefs. They kept a record of their highs and lows in this spreadsheet.

Layer separation: Record of efforts deployed to produce a layer separation model in Rodan that worked across multiple Bach manuscripts. (Spoiler alert, this failed.)

NEW PLAN!: A decision was made to make a separate layer separation model for each manuscript. The Bach projet ended soon after, which is why there's little info in here.

IC nonsense: Gen started trying to teach the IC to recognize clefs. The Bach projet ended soon after, which is why there's little info in here.

CantusDB Volpiano discrepancies

Salzinnes https://docs.google.com/spreadsheets/d/133lVOVM15l7a6bQKIQije4op363ragG4OXewYop_cbQ/edit?pli=1&gid=0#gid=0

Einsiedeln https://docs.google.com/spreadsheets/d/133lVOVM15l7a6bQKIQije4op363ragG4OXewYop_cbQ/edit?pli=1&gid=0#gid=0

When correction for Salzinnes and Einsiedeln was underway, we noticed that the syllabification logic of CantusDB wasn't always syllabifying words properly. Gen started checking every folio agains the CantusDB manuscript text, and kept track of all mis-syllabified words in these spreadsheets. She also logged the typos she occasionally found. She was later given Debra status and started making the corrections. Later, Dylan implemented a new and improved syllabification logic that fixed everything, so Gen stopped making her manual corrections. IMPORTANT: I (Gen) think that there's still a problem with Saved Syllabified Text interfering with Dylan's syllabification revamp. Link issue.

Cantus ultimus testing

https://docs.google.com/spreadsheets/d/1bNaDkkweEpffpzSH2GkyIN5P_xWNaBqZ/edit?gid=1700575581#gid=1700575581

Some time in 2021 or 2022, Gen tested manuscript mapping onto Cantus Ultimus for Nestor. She kept track of her results in this spreadsheet, but did not indicate what the various colours mean. So who knows what went on.

Great IIIF hunt

https://docs.google.com/spreadsheets/d/1OuyW0aJeoUBFH8yKstv58mekZpUdGQAwmrK6Xbp9xLw/edit?pli=1&gid=2036067854#gid=2036067854

Some time in 2021 or 2022, someone (Alessandra or Nestor are best guesses) asked Gen to make a record of which manuscripts on CantusDB have a IIIF manifest and to try and find some for those that didn't. Colours are explained in this one! But no idea if these results were used for anything.

About

A repo for the many beautiful spreadsheets created over the course of work with DDMAL

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors