A resume parser is software that reads a CV file, such as a PDF, a Word document or a photo, and turns its free text into named, structured fields: the candidate's name and contact details, each job with its employer and dates, education, skills and languages. Recruiting software then saves those fields in candidate records, so recruiters can search, filter and compare people without retyping every CV. A CV parser is the same thing under its British name.
This guide explains what a resume parser does with a real example, how resume parsing works step by step, why parsers sometimes get a CV wrong, and what to look at when you choose resume parsing software or an API. If you just want to see one in action, upload a CV to the free online resume parser and compare the result with the file.
What does a resume parser do?
A resume is written for humans, not for software. Every candidate picks a different layout, writes dates differently and names sections in their own way. Software cannot filter on that. A resume parser bridges the gap: it takes the document in and returns the same set of fields for every CV, whatever the template.
Here is a short resume as a recruiter sees it:
JORDAN LEE
Data Analyst · Austin, TX
[email protected] · +1 512 555 0142
EXPERIENCE
Data Analyst, Brightline Health Mar 2021 – Present
- Built weekly retention dashboards in Tableau and SQL
Junior Analyst, Corvo Retail Jun 2019 – Feb 2021
EDUCATION
B.S. Statistics, University of Texas at Austin, 2019
SKILLS SQL, Python, Tableau, stakeholder communication
LANGUAGES English (native), Spanish (professional)
And here is an excerpt of what a resume parser returns for it. This one is shaped like the response of the HireLayer resume parsing API, trimmed for readability:
{
"info_candidate": {
"full_name": "Jordan Lee",
"email": "[email protected]",
"phone_number": "+15125550142",
"job_title": "Data Analyst",
"experience_level": "5 to 10 years",
"location": { "city": "Austin", "region": "Texas", "country_code": "US" }
},
"work_experiences": [
{
"company_name": "Brightline Health",
"job_title": "Data Analyst",
"start_date": "2021-03-01",
"currently_active": true
},
{
"company_name": "Corvo Retail",
"job_title": "Junior Analyst",
"start_date": "2019-06-01",
"end_date": "2021-02-01",
"currently_active": false
}
],
"educations": [
{
"degree_title": "B.S. Statistics",
"school_name": "University of Texas at Austin",
"end_date": "2019-01-01"
}
],
"skills": [
{ "skill_title": "SQL", "skill_type": "Software skill" },
{ "skill_title": "Tableau", "skill_type": "Software skill" },
{ "skill_title": "Stakeholder communication", "skill_type": "Soft skill" }
],
"languages": [
{ "language": "English", "level": "Native or Bilingual Proficiency (C2)" },
{ "language": "Spanish", "level": "Professional Working Proficiency (B2)" }
]
}
Three things happened that matter more than the extraction itself:
- Dates became comparable values. “Mar 2021 – Present” is now a start date and a current-role flag, so software can compute tenure or sort by most recent job.
- Free text became categories. “Spanish (professional)” is now a level on the CEFR scale, and the total experience falls into a bracket that a recruiter can filter on.
- Skills became separate items. A comma-separated line is now a list where each skill has a type, ready for search or for matching against a job.
How resume parsing works, step by step
Every resume parser, whatever its technology, has to solve the same chain of problems. Research on resume information extraction has described the task this way for two decades: first split the document into labeled blocks, then extract detailed fields inside each block.
1. Get the text out of the file
A PDF exported from Word or Google Docs carries a text layer that can be read directly. A scanned page or a phone photo has no text at all: the characters must first be recognized from the image, which is called optical character recognition (OCR). OCR quality depends on resolution, contrast and skew, which is why scans parse less reliably than native files.
2. Rebuild the reading order
Text inside a file isn't stored in the order a person would read it. Two-column templates, sidebars, tables and text boxes all have to be put back in a sensible sequence. When this step fails, a job title from the left column can end up glued to a skill from the right column.
3. Find the sections
The parser identifies which lines belong to contact details, work experience, education, skills, languages or certifications. Standard headings such as “Experience” help, but a good parser also recongizes sections from their content when the heading is creative or missing.
4. Extract the fields
Inside each section, the parser pulls out individual values: employer, job title, start and end dates, degree, school, each skill. It also has to decide where one job ends and the next one begins, wich is harder than it looks when a candidate held two roles at the same company.
5. Normalize the values
Raw strings are converted into consistent formats: ISO dates, country codes, language levels, education levels, contract types, deduplicated skill names. Normalization is what makes the ouput usable in filters and reports, and it is where parsers differ most.
6. Check the document and return the result
Finaly, the output is returned as structured data, usually JSON or XML. A
careful parser also rejects inputs that are not resumes, such as a cover
letter, an ID card or an empty scan, instead of creating an empty candidate
profile. The HireLayer API, for example, answers HTTP 422 with codes such as
DOCUMENT_NOT_A_RESUME or DOCUMENT_TEXT_EMPTY in
those cases.
What information does a resume parser extract?
The exact list depends on the product, but most resume parsers cover the groups below. Check the field list of any parser against what your own software needs to store.
| Field group | Typical fields | What it is used for |
|---|---|---|
| Contact details | Name, email, phone number, LinkedIn and other profile URLs | Creating the record, deduplicating, reaching out |
| Location | City, region, country, postal code, willingness to relocate | Searching by area and commute |
| Work experience | Employer, job title, dates, current role, description, contract type | Tenure, seniority, relevant past roles |
| Education | Degree, field, school, dates, education level | Minimum qualification filters |
| Skills | Individual skills with a type (technical, software, soft skill) | Keyword search and candidate matching |
| Languages | Language and level, often mapped to CEFR | Language requirements |
| Other | Certifications, interests, driving licences, availability | Role-specific filters |
HireLayer's field-by-field reference is in the CV Extract documentation, with types and example values for each field.
Types of resume parsers
Resume parsing technology has gone through three broad generations. Many products today mix them, so treat these as families rather than strict boxes.
- Keyword and rule-based parsers look for known headings, patterns such as email addresses and date formats, and lists of job titles or skills. They are fast and predictable, but break on layouts and wordings their rules did not anticipate.
- Statistical and machine learning parsers are trained on labeled resumes to recognize sections and entities from context. They handle variety better, but need large training sets for each language and resume style.
- Parsers built on large language models read the resume more like a person would and cope well with unusual layouts and wording. The challenge moves to consistency: returning the same normalized fields every time, at a predictable cost and latency.
If you're buying, the technology label matters a lot less than the observable result on your own CVs: which fields come back, how often they are right, how values are normalized and how long a parse takes. Our guide to resume parsing accuracy shows how to measure that field by field.
Who uses resume parsers?
- Applicant tracking systems parse every application so recruiters can search the candidate database. See how an ATS resume parsing API fits into intake.
- Job boards and career sites prefill the application form from an uploaded CV, so candidates confirm their details instead of typing them.
- Staffing and recruitment agencies turn inboxes full of attachments into a searchable talent pool, often in bulk.
- HR tech and AI products use parsed fields as the input for candidate matching, ranking or analytics.
- Job seekers run their own CV through a parser to check that software reads it correctly before applying.
Resume parser vs. ATS vs. resume screening
These terms are often mixed up. A resume parser only turns a document into data; it does not judge the candidate. An applicant tracking system is the larger product that stores candidates, manages job openings and moves people through hiring stages; most ATSs include or license a parser. Resume screening or candidate matching compares parsed profiles with a job's requirements and scores or ranks them, which is a seperate step with its own rules. HireLayer offers it through separate candidate matching and candidate ranking APIs, whose scores support a recruiter's decision rather than replace it.
The distinction matters for compliance too. Under the EU AI Act, scoring or ranking candidates is listed as high-risk, while the Commission's draft guidelines let a system that only organizes CV information into a searchable database use an exception, provided the assessment is documented. Our EU AI Act recruitment checklist covers what that means in practice.
Why resume parsers get CVs wrong
In practice, most parsing errors come from the file itself rather than from what's written in it. The usual causes are:
- Multi-column layouts and sidebars that mix up reading order.
- Text inside images, logos, headers or text boxes, which may be skipped entirely.
- Scans and photos, where every character has to be recognized first.
- Inconsistent or missing dates, which make it hard to split and order jobs.
- Skills shown as rating bars or icons, which contain no readable text.
- Creative section titles such as “My journey” instead of “Experience”.
For software teams, the right response is to keep uncertain or missing values reviewable instead of treating an empty field as proof that the candidate lacks that information. For candidates, a simple one-column file with real text and consistent dates avoids almost all of these problems.
How to check how a parser reads your resume
Honestly, the quickest test is to run the file through a parser and look at what comes back. Upload it to the free online resume parser to see every extracted field and the raw JSON, or use the ATS resume checker if you want a readability score and a list of fixes. Neither requires an account. Keep in mind that every applicant tracking system uses its own parser and settings, so results can differ from one system to another; a file that parses cleanly on one usually follows the basics that help everywhere.
A low-tech check also works: open your PDF, select all the text and paste it into a plain text editor. If the order is scrambled or whole sections are missing, a parser will likely struggle too.
How to choose resume parsing software or an API
If you are adding resume parsing to a product, compare vendors on criteria you can verify, ideally with a trial on a few hundred of your own CVs:
- Formats: PDF and DOCX are a given; check older Word files, OpenDocument, RTF, plain text and images if your candidates send them.
- Languages: test every language in your applicant mix, not only English.
- Fields and normalization: are dates, levels, countries and skills returned as consistent values you can filter on, or as raw strings you must clean yourself?
- Errors: does the API tell you clearly when a file is not a resume, is unreadable or is empty?
- Latency and volume: synchronous responses suit upload forms; bulk imports need throughput and predictable rate limits.
- Pricing: per parse, per seat or by subscription tier, and whether failed calls are billed. See our breakdown of resume parser API pricing.
- Data handling: where files are processed, how long they are kept and whether you can opt out of storage. Resumes are personal data, so GDPR and similar laws apply.
For a side-by-side view of public specifications and a test plan you can reuse, read the resume parser API comparison, or see how HireLayer compares with Affinda, Textkernel and other resume parsers.
See a resume parser on your own CV
HireLayer CV Extract reads 13 file formats and resumes in 70 languages and returns the same JSON fields for every CV. Try it without an account, then call the API with 50 free credits a month.
Frequently asked questions
Is a CV parser the same as a resume parser?
Yes. “CV” is the usual term in the UK, Europe and much of the world, “resume” in North America. CV parsing and resume parsing describe the same process and the same software.
How accurate are resume parsers?
It depends on the field, the file and the parser. Contact details are usually read very reliably, while job boundaries, dates and skills vary more, especially on scans and multi-column templates. Vendor-wide percentages say little about your documents: measure precision and recall per field on a sample of your own CVs.
Can I parse a resume for free?
Yes. The HireLayer free online resume parser handles one CV at a time with no sign-up, and the API's free plan includes 50 credits a month for developers. Open-source libraries exist too, but you host, tune and maintain them yourself.
Does an ATS reject resumes it cannot parse?
Parsing itself does not accept or reject anyone: it fills fields. But a field the parser missed, such as a job title or a skill, can make a candidate harder to find in searches and filters. That is why a readable file matters even when a person makes the final decision.
What file format is best for resume parsing?
A PDF exported from a word processor, or a DOCX file, with real text in a single column. Avoid scans, photos and designs where text sits inside images. If you can select a sentence in your PDF viewer, a parser can read it too.
Do resume parsers work in languages other than English?
Good ones do, but coverage varies by vendor and by field. HireLayer supports resumes in 70 languages and returns the same field names whatever the language of the CV. Always test the languages your candidates actually use.
Sources and further reading
- HireLayer API documentation: CV Extract fields and errors
- Yu, Guan and Zhou (ACL 2005): Resume Information Extraction with Cascaded Hybrid Model
- Tesseract documentation: Improving the quality of the output
- Council of Europe: Common European Framework of Reference for Languages (CEFR)
- JSON Resume: an open schema for resume data
Louis Desclous
Published on · Reading time: 10 minutes

