kobimusic

workspace

← learn

for institutions: big jobs

reading a whole collection: which way to go, how to prepare it, and what to expect

This is for libraries, archives, universities and publishers with a collection to read, from a few hundred pages to many thousands. It is the big picture. The other pages have the details.

which way to go

  • The workspace, for a few hundred pages. Put each piece in a folder, right-click the folder, and choose “document to sheet music”. There is nothing to set up. Getting started has the tour.
  • The API or an agent, for scripted batches. Make an API key under “API keys” on the settings page, or connect an agent over MCP, and send one piece at a time from a script. See the API quick start and the agents quick start.
  • Enterprise, for a whole collection. We host the job, and send you back the outputs. It is priced from 2¢ a page with copista-mini. Ask for a quote, and see “asking for a quote” below.

The API spends the same credits as the workspace, so it does not raise your allowance. What it gives is a script that can run unattended. A run in the workspace goes on only while its page is open.

Credits set the pace. A page of sheet music is one credit, and Personal ($25 a month) has 200 credits a day, so a thousand pages take 5 days and ten thousand take 50 days. The free plan's 30 a week is for trying it on a few of your own pages. When the days add up to more than you want to wait, or you would rather not run it yourselves, ask for a quote.

the limits that matter at scale

limit
credits30 credits a week on the free plan, 200 credits a day on Personal; a page of sheet music is one
pages in one piece50
one file, and one call95 MB
files in one run2,000; a folder stands for every file in it, subfolders too
room for your files512 MB, and 2,000 files and folders
calls at once2 for one person
through the APIuploads and results are kept for a quarter of an hour

Limits has every cap. Two of these decide how a big job is run. Credits come in a day's or a week's worth at a time, and a run that needs more than is left is not started, so choose a day's worth of folders and the menu will say what it costs before you begin. And 512 MB of room is soon full of scans, so work in batches: bring in a folder, run it, download the results, delete what you no longer need, and bring in the next.

preparing a collection

  • One piece, one PDF, or one folder of page images. A PDF is always read as one piece. The images in a folder are read as one piece too, in the order of their names, and you can say instead that every image is its own piece, for a folder of single pages.
  • At most 50 pages to a piece. A longer volume is turned away, so split it first into parts of up to 50 pages and read each as its own piece.
  • Name page images so they sort in reading order. “page 2” comes before “page 10”, and zero-padded numbers such as page-001.png are the safest.
  • Use PDF, PNG, JPEG or WebP. HEIC photos are not read yet, so export them as JPEG. A multi-page TIFF has only its first page read, so make it a PDF, or one image to a page.
  • A part-book is a document where each instrument has its own part, one after another, such as the four parts of a string quartet. Say so when you run it, so the parts are read to play together and not as one long score. In the workspace, tick “a part-book” when it asks. In the API, send partBook=true. In a flow, it is the partBook parameter. A run asks once for everything you chose, so keep part-books in folders of their own and run them apart from the rest.

Scans that help: pages that are straight, not turned or skewed; the whole page in the frame, margins included; and enough resolution that the smallest marks, such as dots, accidentals and small print, are still sharp. A photograph works if the page is flat and evenly lit.

The reader is strongest on real scans, including authentic 1800s and 1900s editions. It works for some handwritten scores, but if even you have trouble reading a page, the model will too. Try a few typical pieces and a few hard ones first, and look at what comes back before you run the rest.

what comes back

Sheet music comes back as compressed MusicXML (.mxl), one file for each piece. sonata.pdf gives sonata.mxl, and a folder of page images gives a file named after what the page names share, or after the folder. MuseScore, Finale, Sibelius and Dorico open it. A MIDI file for each score is one more step and costs no credits: right-click the .mxl, or a whole folder of them, and choose “sheet music to midi”.

In the workspace, results are saved beside what they came from, so a folder of pieces comes back as the same folders with a score beside each. To take many away, right-click the folder, choose “compress into a .zip”, and download the .zip. “download” is for one file at a time. A .zip can hold up to 128 MB and 2,000 files, and it is a file in your files too, so delete it once you have it.

Through the API, the file is the body of the answer. What the reader had to say about the call, such as how many pages it read, the credits it cost and which model ran, is in the X-Kobi-Report header, or in the JSON answer. Keep it beside the score if you want a record.

checking the results

  • Nobody checks the results for you. They come straight from the model, and may be incomplete or wrong, so plan your own review before anything is published or performed from them.
  • Spot-check as you go. Open a score in the preview, which plays it, and compare a few pages with the scan. Look at the first pieces from each kind of source before you run the rest.
  • The tasks page lists the tasks of a run, one for each piece, and every one that failed, with the reason. A long run's list is trimmed to 1,000 tasks and the failures are kept first. The last 200 runs are kept, for 30 days.
  • A page with no music on it comes back as “no music was found”, and costs nothing. A file the reader cannot open at all costs one credit.

your files and your data

Files in your account are encrypted on the server and kept until you delete them. Deleted files are removed within three days. The server can still read them, since it needs to in order to run the tools on them.

What you upload, and what the tools make from it, may be used to train and evaluate kobimusic's models. The terms and the privacy policy say how: it is not published or made available to others, and deleting a file ends the use for the future but does not undo what was already used. They also ask you to have the rights to what you upload, including the rights in an edition as well as in the music. If your collection needs different terms, say so when you ask for a quote. A quote is a written agreement, and it governs where it differs from the terms.

asking for a quote

Ask for a quote with what you have: what the files are, about how many pages or hours, and what you want back. Each quote says what a page is, the formats, and when it will be done. We host the job and send you back the outputs, which come straight from the models, and nobody corrects them.

To see the quality first, make a free account and read a few of your own pages: 30 credits a week is 30 pages.