EML: UTF-8 Filename Attachment
An attachment whose filename uses RFC 2231 UTF-8 encoding (résumé), for MIME filename decoders.
Content-Type: multipart/mixed; boundary="=_fn_utf8_0001"
MIME-Version: 1.0
From: Docs <docs@meridiansupply.example>
To: Sam Rivera <sam@meridiansupply.example>
Subject: =?utf-8?b?QXR0YWNoZWQgcsOpc3Vtw6k=?=
Date: Sun, 11 Jan 2026 12:00:00 -0800
Message-ID: <fn-utf8@meridiansupply.example>
--=_fn_utf8_0001
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: base64
UGxlYXNlIGZpbmQgdGhlIFNBTVBMRSByw6lzdW3DqSBhdHRhY2hlZC4K
--=_fn_utf8_0001
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: base64
Content-Disposition: attachment; filename*=utf-8''r%C3%A9sum%C3%A9-sample.txt
QWRhIExvdmVsYWNlIOKAlCBTQU1QTEUgcsOpc3Vtw6kKRXhwZXJpZW5jZTogYW5hbHl0aWNhbCBl
bmdpbmUuCg==
--=_fn_utf8_0001--
Specifications
- Filename Encoding
- RFC 2231 utf-8
- Filename
- résumé-sample.txt
Testing contract
Expected to pass- Scenario
- Exercise EML: UTF-8 Filename Attachment in its mime workflow. An attachment whose filename uses RFC 2231 UTF-8 encoding (résumé), for MIME filename decoders.
- Expected result
- subject='Attached résumé'; content type=multipart/mixed; MIME leaf parts=2. Declared feature checks: filenameEncoding=RFC 2231 utf-8; filename=résumé-sample.txt.
What is a .eml file?
An EML file is a single email message stored in the RFC 822 / MIME format: plain-text headers (From, To, Subject, Date, Message-ID) followed by the body, which may be plain text, HTML, or a multipart structure with alternative bodies and file attachments encoded in base64.
How to use this file
Use an example EML to test email header parsing, MIME decoding, HTML-part handling, attachment extraction, and EML-to-other-format conversion.
How to use this file for testing
“EML: UTF-8 Filename Attachment” is a deterministic Testaroo fixture for Email parsing, Encoding detection, Internationalization. Standards-compliant RFC 822 messages (plain, multipart text+HTML, and with an attachment) plus an MBOX mailbox, for testing header parsing, MIME decoding, attachment extraction, and mailbox splitting.
Documented properties for this file: EML · 760 bytes. Compare results against paired or grouped companions on this page when present (clean↔damaged, searchable↔scanned, or format twins) so scores stay reproducible across runs.
Download the file once, keep the path stable in CI or local scripts, and treat the spec table as the contract: dimensions, seeds, field lists, and roles are intentional. Corrupt or invalid samples are labelled as such, expect parsers to fail loudly rather than silently accept them.
Email fixtures use fixed dates, message IDs, and MIME boundaries so runs are reproducible, and every address is fictional. Test header parsing, MIME decoding, attachment extraction, and EML/MBOX conversion against the documented structure.
Code examples
from email import policy
from email.parser import BytesParser
msg = BytesParser(policy=policy.default).parse(open("utf8-filename-attachment.eml", "rb"))
print(msg["subject"], msg["from"])Generated by generation/email_wave_c.py. Free for any use, no attribution required, license.
Related files
- emlEML: RFC 2231 Continuation Filename With size and creation-dateThe conformant answer to non-ASCII filenames: a filename split into three numbered RFC 2231 segments with a charset and a language tag, plus size and creation-date parameters. Parsers commonly handle filename*= but not the numbered continuation form.

- emlEML, RFC 2047 Subject: ISO-8859-1 'B' EncodingThe ISO-8859-1 Subject again, base64-encoded instead of Q-encoded. Base64 hides the charset entirely, so this is the fixture that catches a decoder guessing the charset from the octets instead of reading the encoded-word's charset token.

- emlEML, RFC 2047 Subject: ISO-8859-1 'Q' EncodingA German Subject encoded as ISO-8859-1 'Q', the shape most legacy mail actually uses. Each umlaut is one =XX escape, so a decoder that assumes UTF-8 produces mojibake rather than a clean error.

- emlEML, RFC 2047 Subject: KOI8-R 'B' EncodingA Russian Subject as a KOI8-R base64 encoded-word. KOI8-R orders Cyrillic letters by Latin transliteration rather than alphabetically, so a decoder that substitutes any other Cyrillic codepage returns readable-looking but wrong text.

- emlEML, RFC 2047 Subject: Shift_JIS 'B' EncodingA Japanese Subject as a Shift_JIS base64 encoded-word. Shift_JIS second bytes overlap ASCII punctuation values, so a decoder that scans the decoded octets for delimiters before converting the charset splits the string in the wrong place.

- emlEML, RFC 2047 Subject: UTF-8 'B' EncodingA Subject header carrying accented Latin text and an emoji as a single UTF-8 base64 encoded-word. Its Q-encoded twin decodes to the identical string, so the two together isolate the encoding from the charset.
