EPUB Files, Converted and Corrected

After reading this you will understand what sits inside an EPUB file, how to pull its chapters out as clean text or Markdown, and how to fix broken title and author fields without breaking the book.

What an EPUB actually is

An EPUB is not a single document. It is a ZIP archive with a fixed internal layout. Rename a file from book.epub to book.zip, unzip it, and you will find a folder tree of files. Open one in a text editor and you are looking at XHTML, the same markup a web page uses.

A minimal EPUB contains three kinds of parts. First, one file called mimetype that holds the exact string application/epub+zip and nothing else. Second, a manifest file (the OPF, usually content.opf) that lists every resource and records the metadata. Third, the content itself: one XHTML file per chapter, plus CSS, fonts and images.

The OPF is the map. It has a <manifest> listing all files and a <spine> listing the reading order by reference. The spine matters. Files inside the ZIP can be stored in any order, so the tool follows the spine rather than the ZIP directory to produce chapters in the order a reader would meet them.

The first entry in an EPUB ZIP must be the uncompressed mimetype file, stored at a known byte offset. That is why a reader can identify an EPUB from its first few dozen bytes without unzipping anything. You can see those bytes with the Hex Viewer.

When to convert, and when not to

Convert to plain text when you want the words and nothing else: to paste a passage into an email, to run a word count, to feed a chapter to a script, or to read on a device that only shows .txt. Convert to Markdown when structure matters: you want headings, bold and italic, links, lists and blockquotes preserved so the text still reads like a book.

Do not expect a faithful visual copy. EPUB layout comes from CSS, and CSS does not survive conversion to text. A two-column sidebar, a drop cap, a decorative pull quote: all of these flatten into ordinary paragraphs. Complex layout tables become linear text, so a 4-column pricing grid may read as a run of cells with no columns.

The tool cannot open DRM-protected books. Adobe Digital Editions files and Kindle formats (AZW, KFX, MOBI) are encrypted or use a different container entirely. Only DRM-free EPUB 2 and EPUB 3 files open here. If a shop lets you download a plain .epub with no login required to read it, that is usually DRM-free.

How the conversion works, step by step

There is no floating-point math here, so the "formula" is a procedure. The tool runs it in this order:

  1. Read the ZIP directory and locate META-INF/container.xml, which names the OPF file.
  2. Parse the OPF. Read the metadata block and the spine list.
  3. For each spine entry in order, load its XHTML file.
  4. Walk the XHTML tree. Keep text nodes. Map structural tags to output: <h1> becomes a Markdown # heading, <strong> becomes **bold**, <ul><li> becomes - lines.
  5. Drop presentational tags (<span>, <div> used only for styling) and collapse the runs of whitespace that XHTML pretty-printing leaves behind.

Plain text output takes the same walk but throws away the markers: a heading is just its words on their own line, emphasis is just the word. That is why Markdown output is always at least as large as text output for the same book.

A worked example with the demo book

The demo button loads a small sample EPUB with default settings. Walk through the numbers so you can check them against your own run.

Following the spine of the sample

Suppose the sample's OPF spine lists five items in this order: cover, toc, ch1, ch2, ch3. Inside the ZIP the files happen to be stored alphabetically as ch1, ch2, ch3, cover, toc. The spine wins, so export produces cover first and chapter 3 last, not the ZIP order.

  1. Chapter 1 XHTML holds one <h1> ("Chapter One"), 3 paragraphs, and one bold phrase.
  2. Markdown output: # Chapter One on line 1, a blank line, then the 3 paragraphs, with the bold phrase wrapped in **. Character count roughly 1,240.
  3. Text output of the same chapter: Chapter One with no #, no **. Character count roughly 1,228, which is 12 characters less: two #-and-space pairs and four asterisks removed.
  4. Export all as ZIP produces 5 files named by spine order with a two-digit prefix: 01-cover.md through 05-ch3.md.

The numbering prefix is what keeps files sorted in reading order once they land in your file manager, which sorts by name.

Markdown is slightly larger than plain text for each chapter because it keeps the heading and emphasis markers. The gap grows with the number of headings and emphasised runs.

Editing metadata without breaking the file

Metadata lives in the OPF as Dublin Core elements. A title is <dc:title>, an author is <dc:creator>, the language is <dc:language>, and the blurb is <dc:description>. Editing these changes what your library app shows in its shelf view. It does not change a single word of the chapters.

The language field expects a BCP 47 code, not a word. Write en for English, en-US for American English, fr for French, de for German, ja for Japanese. A reader that sees english instead of en may fall back to a default and hyphenate or sort your book wrong.

Saving a corrected EPUB rewrites the OPF and repacks the ZIP. That means the new file's bytes differ from the original, so its checksum changes even though the reading text is identical. Do not be surprised when the File Hash Checker reports a different SHA-256 for the saved copy.

Fixing a swapped title and author

A common import bug leaves a book shelved with author "Untitled" and title "Jane Austen". Open the file, read the current values, put Pride and Prejudice in the title field and Jane Austen in the author field, set language to en, and save. The tool writes:

<dc:title>Pride and Prejudice</dc:title>
<dc:creator>Jane Austen</dc:creator>
<dc:language>en</dc:language>

Delete the old file from your library and re-import the saved one. The shelf now sorts under P for the title and A for the author.

See the effect of format choice

The one thing a static example cannot show is how the same source markup lands in each output. Move the toggle and watch the same chapter change shape.

Given a fixed sample chapter with one H1 heading, one bold phrase, one link and a three-item bulleted list: the Markdown output keeps # Heading, **bold**, [text](url) and - item lines; the plain-text output shows the heading as bare words, the bold phrase as plain words, the link as its visible text only, and the list as three lines with no dashes. Markdown here is 41 characters longer than the text version.

Common mistakes

Expecting page numbers. EPUB is reflowable text with no fixed pages, so there is nothing to convert into "page 42". A citation that needs a page number needs the print edition or a PDF.

Assuming chapter files match the table of contents. The visible table of contents (the NAV document) and the spine are separate lists. A book can have 30 spine files but a TOC that only names 12 of them, because front matter and section dividers sit in the spine without a TOC entry. Export follows the spine, so you get all 30.

Treating a corrupt ZIP as a DRM problem. If the file will not open, check whether it is a valid archive first. Try unpacking it with the ZIP Creator & Extractor. If that fails too, the file is damaged, not encrypted, and no ebook tool will read it.

Forgetting encoding. XHTML in EPUB is almost always UTF-8, and output here is UTF-8. If you open the exported .txt in an old editor set to a legacy code page, accented letters and curly quotes will show as garbage. The bytes are correct; the editor is guessing wrong.

Related tools on this site

An EPUB is a ZIP, so the ZIP Creator & Extractor can open one directly if you want to poke at the raw XHTML and OPF yourself. Use the File Hash Checker to confirm a download matches a published checksum before you trust it, and to see how repacking changes the hash. When a file refuses to open and you want to know whether it even starts with the EPUB signature, the Hex Viewer shows you the first bytes.

Frequently asked questions

Does anything get uploaded when I convert a book?

No. The whole process runs in your browser. The file is read, unzipped and converted on your device, and nothing is sent to a server. You can turn off your network connection after the page loads and it still works.

Can it open Kindle books or DRM-protected EPUBs?

No. Kindle formats (AZW, KFX, MOBI) use a different container, and DRM-protected files are encrypted. Only DRM-free EPUB 2 and EPUB 3 files open. A store download that requires special reader software to view is almost always DRM-protected.

Why are chapters in a different order than the table of contents?

Chapters follow the spine, the OPF's reading-order list. The visible table of contents is a separate list and can skip files. If the two disagree, the spine is the true reading order.

Will Markdown keep images and tables?

Images are not embedded in the text output. Simple structure (headings, emphasis, links, lists, blockquotes) is kept in Markdown. Layout tables are flattened to linear text, so column alignment is lost.

Why does the saved EPUB have a different file hash?

Saving rewrites the OPF and repacks the ZIP, so the byte sequence changes even when the reading text is identical. A changed hash is expected and does not mean the content was altered.