Glossary term
CV Parsing: Definition and How It Differs From Formatting
CV parsing is the automatic extraction of a CV's contents into structured data. Learn how it works, where it goes wrong and how it differs from formatting.
CV parsing, also called resume parsing, is the automatic extraction of the information in a CV, such as name, contact details, employers, job titles, dates, education and skills, into structured data fields that software can store, search and reuse.
Parsing is what happens when you drag a CV into an applicant tracking system or recruitment CRM and the candidate record fills itself in. It is also the first step inside most CV formatting tools.
How CV parsing works
- Text extraction. The software reads the text from a PDF or Word file. Files that are scanned images need optical character recognition (OCR) first, which adds another source of error.
- Structure detection. It identifies sections such as experience and education, and works out which lines belong together.
- Field extraction. It assigns values to fields: this line is an employer, that one a job title, these are start and end dates.
- Normalisation. Some parsers standardise values, for example converting "Sept 21" into a date or mapping a skill to a taxonomy.
Older parsers relied mainly on rules and statistical models. Many current tools use large language models, which cope better with unusual layouts but can introduce new kinds of mistake.
Where parsing goes wrong
- Complex layouts. Two-column designs, tables and text boxes can scramble the reading order.
- Lost qualifiers. "MSc (not completed)" can become a completed MSc.
- Date errors. A missing end date can become "Present", and overlapping roles can be merged.
- Misattribution. A location or skill mentioned once can be attached to the wrong role.
- Over-normalisation. "Basic conversational Spanish" mapped to a formal proficiency level the candidate never claimed.
None of these is unusual, which is why parsed data should be checked before it is relied on, especially when it will be sent to a client. Our article on AI CV formatting accuracy explains how to test for them.
Parsing versus formatting
Parsing produces data: fields in a database. Formatting produces a document: a CV laid out in an agency or client template. Formatting tools usually parse first, then render the parsed data into a template. The difference matters when choosing software. A parser feeding your ATS needs good search fields; a formatter needs to preserve every fact exactly and produce a clean, editable document. Read CV parsing vs CV formatting for a fuller comparison, and see the CV formatting entry.
How CVPitch uses parsing
CVPitch reads PDF and Word files directly, converts older .doc, .rtf and .odt files, and reads scanned PDFs and photos with AI text recognition (OCR) at 1 AI credit per page, flagging the transcription for review. It then uses Anthropic's Claude (Haiku) to extract facts into structured fields, with a hard cost cap per CV. Each extracted fact must point to the exact passage in the source; facts without supporting evidence are marked unverified, and uncertain items become review flags. The candidate document is treated as data, never as instructions to the AI. A recruiter reviews and approves before anything is exported. See features and security.
Related terms
- CV formatting: turning parsed data into a client-ready document.
- Applicant tracking system: where most parsed CVs end up.
- Recruitment CRM: the relationship system that stores candidate records.
Want to see it on your own CVs? Start a free trial with 10 exports and no card, compare plans on the pricing page, or browse all features and integrations.
Turn your next CV into a client-ready submission
Upload a candidate CV, check every fact against the source, remove contact details and export your agency’s branded DOCX and PDF. Your first 10 CVs are free.