How the ATS simulator actually works
The rubric behind every score, the step-by-step algorithm for each of the four assessment stages, which API is called at each point, and the real code for the parts we built ourselves.
Read this first. ValidateCV is a simulator, not a copy of any specific employer's ATS. Vendors such as Workday, Greenhouse and Lever do not publish their parsing rules. Our rubric encodes widely documented ATS failure modes. We mark every heuristic as a heuristic and list what is not finished yet in Known limitations.
The pipeline at a glance
One scan is a single run through the steps below. Green panels run entirely in your browser. The blue panel is the only place text leaves your device, and by then it has been redacted. The dashed panel is the paid stage.
Your device
Browser only. Nothing here touches the network.
Add resume and job description
Drag-and-drop PDF or DOCX, plus pasted text or an uploaded file for the job.
File API
Extract text from the files
PDF text layer via a Web Worker; DOCX raw text. 20-second timeout guard.
pdfjs-dist (legacy) / mammoth
Redact personal details
Emails, addresses, phone numbers and the name line are replaced with tokens.
lib/redact-pii.ts (regex)
Hold the redacted text for this tab
Saved so the report page can read it. Cleared when the tab closes.
sessionStorage
ValidateCV server and AI model
Receives redacted text only. Stores nothing.
Receive redacted text
Validates the request body with a schema, then forwards it for extraction.
POST /api/extract (Next.js Route Handler)
Turn text into structured data
Two parallel model calls return typed JSON: roles, skills and dates from the resume; must-haves and knockouts from the job.
AI SDK generateObject + zod, via Vercel AI Gateway (openai/gpt-4o-mini, temperature 0)
Your device again: free report
Structured JSON comes back. Scoring runs locally, with no more AI calls.
Stage 1: Parsing and readability
Five formatting checks on the redacted text, averaged into a score.
analyzeParsing()
Stage 2: Experience and recency
Years, current role and employment-gap detection from the extracted roles.
analyzeExperience()
Stage 3: Keyword and knockout match
Each job requirement is searched for in the resume corpus.
matchRequirements()
Stage 4: Deep semantic match
Endpoint built and tested. Shown as a locked preview in the report today.
Judge each requirement 0 to 3
The model scores and quotes evidence. A TypeScript function does the weighted maths.
POST /api/assess, generateObject via AI Gateway (temperature 0)
The ATS rubric
The report has four stages. Stages 1 to 3 are the free Guest Check: they flag red flags and never call an AI model to score you. Stage 4 is the only stage where a model forms a judgement.
| Stage | What it checks | Method | How it is scored | Tier |
|---|---|---|---|---|
| 1. Parsing and readability | Whether a parser can read your layout: length, columns, headers, tables, special characters. | Deterministic code | Five checks, each pass (100), warning (60) or fail (20). Score is the mean. | Free |
| 2. Experience and recency | Total years, most recent role, and unexplained employment gaps. | Model extracts dates; code analyses them | Informational. Gaps are advisory and never lower a score. | Free |
| 3. Keyword and knockout match | Whether each must-have and nice-to-have requirement appears in your resume. Lists knockout criteria. | Model extracts requirements; code matches them | Matched count per group. Knockouts are shown, not scored. | Free |
| 4. Deep semantic match | Whether your experience genuinely satisfies each requirement, beyond exact words. | Model judges; code calculates | Each requirement scored 0 to 3, weighted, converted to a 0 to 100 score. | Paid |
Design principles behind the rubric
- Models extract and judge; code scores. No percentage you see is invented by a language model. Models produce structured facts or a 0 to 3 grade, and plain TypeScript turns those into numbers.
- Temperature 0 everywhere. Extraction and judging are classification tasks, so the same input should give the same output.
- Red flags, not verdicts. Stages 1 to 3 tell you what a bot could trip on. They do not predict whether you will be hired.
Intake: parsing and PII redaction
Your files become plain text on your device, and personal details are stripped before anything is stored or sent.
- Runs on
- Your device (browser)
- Method
- Libraries and regular expressions
- Available to
- Everyone, including guests
APIs and libraries used in this step
File.arrayBuffer()
Browser API · reads the upload; the file is never sent anywhere
pdfjs-dist (legacy build)
Library (in browser) · PDF text layer, runs in a Web Worker
mammoth
Library (in browser) · DOCX raw text
lib/redact-pii.ts
Our code · section-aware, one-way redaction
sessionStorage
Browser API · redacted text only, per tab
Algorithm
- Detect the file type from the MIME type or the extension, then route to the PDF or DOCX parser.
- Extract text locally. For PDFs, group text fragments into real lines by their coordinates, then drop page numbers and any top or bottom line that repeats across pages (running headers and footers). If no text exists (a scanned image), the run continues with a placeholder and you are warned.
- Race the parser against a 20-second timer. A stalled PDF worker can never freeze the page.
- Run
redactResume()on the resume and the lighterredactPII()on the job description. - Write only the redacted text to sessionStorage, then navigate to the report.
Novel solution 1: a parser that cannot hang
Chrome stalled on the current PDF.js build because its worker relies on JavaScript APIs Chrome does not ship yet. We import the legacy build, which bundles polyfills, and we wrap every parse in a timeout race.
// lib/parse-file.ts (runs in the browser)
const PARSE_TIMEOUT_MS = 20_000
export async function parsePdfToText(file: File): Promise<string> {
// The legacy build ships polyfills Chrome may lack. The modern build
// crashes its worker there and getDocument() never settles.
const pdfjsLib = await import('pdfjs-dist/legacy/build/pdf.mjs')
pdfjsLib.GlobalWorkerOptions.workerSrc = new URL(
'pdfjs-dist/legacy/build/pdf.worker.mjs',
import.meta.url,
).toString()
const doc = await pdfjsLib.getDocument({ data: await file.arrayBuffer() }).promise
const pageTexts: string[] = []
for (let n = 1; n <= doc.numPages; n++) {
const content = await (await doc.getPage(n)).getTextContent()
pageTexts.push(content.items.map((i) => ('str' in i ? i.str : '')).join(' '))
}
return pageTexts.join('\n\n').replace(/[ \t]+/g, ' ').trim()
}
export async function parseResumeFile(file: File): Promise<string> {
const isPdf = file.type === 'application/pdf' || file.name.toLowerCase().endsWith('.pdf')
const parsing = isPdf ? parsePdfToText(file) : parseDocxToText(file)
// A worker that fails to start can leave parsing pending forever.
// Race it against a timer so the UI can never freeze.
let timeoutId: ReturnType<typeof setTimeout>
const timeout = new Promise<never>((_, reject) => {
timeoutId = setTimeout(() => reject(new Error('Reading your file took too long.')), PARSE_TIMEOUT_MS)
})
try {
return await Promise.race([parsing, timeout])
} finally {
clearTimeout(timeoutId!)
}
}Novel solution 2: structure-aware, one-way redaction
Redaction happens before any network request, so the server and the model never receive the original identifiers. The resume is split into sections first, because what counts as personal depends on where it sits. The rules run in this order:
- Everything before the first section heading (name, title, subtitle, contact block) collapses into one header token.
- References, referees and personal-details or declaration sections are removed whole.
- In work experience, only the job title, the dates and the task text survive. Employer names, locations and contacts are redacted, and when a line cannot be positively identified as a title or a date it is treated as an employer.
- Employer names (including short forms such as "Acme") and the candidate's name are then scrubbed from the rest of the document, so they cannot reappear in a task bullet or footer.
- Emails, phone numbers, links and handles, street addresses, suburb and state, and labelled personal details (date of birth, gender, marital status, nationality, visa, licence, TFN) are redacted.
- Leftover header and footer lines such as "Name | Page 2" are removed.
Redaction is irreversible from the output alone. Matches record a type and nothing else, every placeholder is a fixed string with no numbering, hash, length or partial digits, a removed block becomes a single token so its size does not leak, and the original text is never written to storage.
// lib/redact-pii.ts (runs in the browser, BEFORE any network call) - excerpt
export function redactResume(text: string): RedactResumeResult {
const lines = text.split('\n').map((line) => line.replace(/\s+/g, ' ').trim())
const headings = lines.map(detectHeading) // "Experience", "References", ...
// 1. Everything before the first section heading is the header block.
const firstSection = headings.findIndex((kind) => kind !== null)
const preSection = lines.slice(0, firstSection)
const nameTerms = namesFrom(preSection)
const out: string[] = []
if (preSection.some(Boolean)) out.push('[REDACTED_HEADER]') // one token: size does not leak
for (let i = firstSection; i < lines.length; i++) {
const heading = headings[i]
if (heading === 'references' || heading === 'personal') {
removedBlock = heading // 2. drop the whole section
out.push('[REDACTED_' + heading.toUpperCase() + ']')
continue
}
if (removedBlock) continue
if (section === 'experience' && entry.mode === 'header') {
// 3. job header lines keep title + dates only
const result = redactJobHeaderLine(lines[i], matches)
companyTerms.push(...result.companyNames.flatMap(companyVariants))
out.push(result.line)
continue
}
out.push(redactEntities(lines[i], matches)) // email, phone, links, address...
}
// 4. Names and employers can reappear in task bullets: scrub them everywhere.
let scrubbed = scrubGlobally(out, nameTerms, 'name', matches)
scrubbed = scrubGlobally(scrubbed, companyTerms, 'company', matches)
// 5. Remove leftover header/footer lines such as "Name | Page 2".
return finish(scrubbed.map((l) => (isFooterLine(l) ? '[REDACTED_FOOTER]' : l)))
}
// "Acme Technologies Pty Ltd" is also written "Acme Technologies" and "Acme".
function companyVariants(name: string): string[] {
const withoutLegal = stripLegalSuffix(name)
const core = stripGenericCompanyWords(withoutLegal)
return [name, withoutLegal, core]
}
// Unsure whether a header segment is a job title? Treat it as an employer.
function classifySegment(text: string): 'title' | 'company' {
const isTitle = !LEGAL_SUFFIX.test(text) && TITLE_WORDS.test(text) && wordCount(text) <= 6
return isTitle ? 'title' : 'company'
}Honest scope of the privacy claim. Redaction is a set of heuristics, not a trained entity recogniser. It does not yet remove university names or the names of people mentioned inside task text, and a resume with no recognisable section headings only loses its leading contact lines and labelled personal details. The redacted text is sent to our server and on to the AI model provider through Vercel AI Gateway. We have no database and our routes do not write your text anywhere, but the provider's own data policy applies to what it receives.
Extraction: unstructured text to structured data
Stages 2 and 3 need facts, such as dates, roles and requirements, not raw paragraphs. One server call turns both documents into typed JSON.
- Runs on
- ValidateCV server, then an AI model
- Method
- LLM extraction at temperature 0
- Available to
- Everyone, including guests
APIs and libraries used in this step
POST /api/extract
Server API · Next.js Route Handler
generateObject()
AI model · AI SDK, schema-validated output
zod
AI model · defines and enforces the output schema
Vercel AI Gateway
AI model · plain model ID, no provider SDK
openai/gpt-4o-mini
AI model · the model used
Algorithm
- The report page reads the redacted text from sessionStorage and POSTs both documents to
/api/extract. - The route validates the body with zod and returns a 400 for anything malformed.
- Two independent
generateObjectcalls run in parallel withPromise.all: one for the resume, one for the job description. - Each response is forced to match its zod schema, so the client always receives well-formed JSON rather than free text.
- The JSON is returned to the browser. Nothing is persisted.
Novel solution 3: schemas as the contract
The model is never asked for a score. It fills a schema, and the field descriptions double as instructions. The job description schema separates must-haves from nice-to-haves and isolates knockouts: hard filters such as work rights or clearance that an ATS auto-rejects on. Prompts are abbreviated below.
// app/api/extract/route.ts (ValidateCV server -> Vercel AI Gateway)
const resumeSchema = z.object({
headline: z.string(),
roles: z.array(z.object({
title: z.string(),
employer: z.string(),
startDate: z.string(), // "2021-03" or "Mar 2021"
endDate: z.string(), // "2023-08" or "Present"
bullets: z.array(z.string()),
})),
skills: z.array(z.string()),
totalYearsExperience: z.number(),
})
const jobDescriptionSchema = z.object({
title: z.string(),
mustHave: z.array(requirementSchema),
niceToHave: z.array(requirementSchema),
knockouts: z.array(z.string()), // hard filters: visa, clearance, mandatory degree
keywords: z.array(z.string()),
})
// The two extractions are independent, so they run in parallel.
const [jdResult, resumeResult] = await Promise.all([
generateObject({
model: 'openai/gpt-4o-mini', // plain Gateway model ID, no provider SDK
schema: jobDescriptionSchema,
temperature: 0, // parsing task: same input -> same output
prompt: '<extraction instructions + redacted job description>',
}),
generateObject({
model: 'openai/gpt-4o-mini',
schema: resumeSchema,
temperature: 0,
prompt: '<extraction instructions + redacted resume text>',
}),
])// app/analysis/page.tsx (the only network call the free report makes)
const response = await fetch('/api/extract', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ resumeText, jobDescriptionText }), // both already redacted
})
const data = await response.json() // { resume, jobDescription } structured JSON
// Everything after this line (Stages 1-3) is computed locally in the browser.Parsing and readability
Can a machine read your resume the way you intended? This stage inspects the extracted text for the layout problems that scramble real parsers.
- Runs on
- Your device (browser)
- Method
- Deterministic code, no AI
- Available to
- Everyone, including guests
APIs and libraries used in this step
analyzeParsing()
Our code · lib/analysis.ts, pure TypeScript
The five checks
| Check | Pass | Warning | Fail |
|---|---|---|---|
| Resume length | 400 to 800 words | 250 to 399 or 801 to 1,000 | Anything else |
| Single-column layout | 15% or fewer of lines look interleaved | More than 15% of lines have 4+ space gaps or exceed 220 characters | Never fails (it is a proxy) |
| Standard section headers | 3 or 4 of Summary, Experience, Education, Skills | 2 recognised | 0 or 1 recognised |
| Tables and text boxes | Up to 5 tab runs or pipe separators | More than 5 | Never fails |
| Fonts and special characters | 20 or fewer decorative glyphs | More than 20 bullets, stars, hearts or emoji | Never fails |
Scoring algorithm
- Run the five checks on the redacted resume text.
- Map each status to points: pass 100, warning 60, fail 20.
- Average the five values and round. Below 70 shows a critical "likely unreadable" flag. 70 to 87 shows a minor-risk notice. 88 and above shows no banner.
Novel solution 4: detecting columns from their symptoms
We receive text, not a rendered page, so we cannot see columns. We detect what columns do to extracted text instead: two columns read line by line produce very long lines and wide internal whitespace. Likewise, fonts are invisible in extracted text, so the font check looks for glyphs known to garble.
// lib/analysis.ts -> analyzeParsing() (pure TypeScript, no network)
const STATUS_SCORE = { pass: 100, warning: 60, fail: 20 }
// Multi-column PDFs extract as interleaved text: long lines or big
// whitespace gaps. We can't see the layout, so we detect its symptoms.
const lines = resumeText.split('\n').filter((line) => line.trim().length > 0)
const suspiciousLines = lines.filter((line) => /\s{4,}/.test(line) || line.length > 220).length
const columnStatus =
lines.length === 0 ? 'warning'
: suspiciousLines / lines.length > 0.15 ? 'warning'
: 'pass'
// Section headers are matched by intent, not exact wording.
const expectedHeaders = [
{ label: 'Summary', pattern: /\b(summary|profile|objective)\b/i },
{ label: 'Experience', pattern: /\b(experience|employment|work history)\b/i },
{ label: 'Education', pattern: /\beducation\b/i },
{ label: 'Skills', pattern: /\bskills\b/i },
]
const found = expectedHeaders.filter(({ pattern }) => pattern.test(resumeText))
const headerStatus = found.length >= 3 ? 'pass' : found.length === 2 ? 'warning' : 'fail'
// Final score: unweighted mean of the five check scores.
const score = Math.round(
checklist.reduce((sum, item) => sum + STATUS_SCORE[item.status], 0) / checklist.length,
)Experience and recency
How much experience you have, what your current role is, and whether there is a timeline gap a recruiter or filter would notice.
- Runs on
- Your device (browser), using extracted data
- Method
- LLM-extracted dates, analysed by code
- Available to
- Everyone, including guests
APIs and libraries used in this step
openai/gpt-4o-mini
AI model · extracts roles and dates in /api/extract
analyzeExperience()
Our code · lib/analysis.ts, pure TypeScript
Algorithm
- Years: taken from the model's
totalYearsExperienceestimate. This number is the model's, not calculated by us. - Current role: the first role whose end date reads "present" or "current", otherwise the first role listed.
- Gap detection: parse each date to a timestamp, compare each role's start with the end of the role before it, and flag the first hole of three months or more.
- Unparseable dates are skipped rather than guessed.
Novel solution 5: advisory gaps
Gaps are real, common and often legitimate, so this check is deliberately incapable of lowering your score. It surfaces an "Advisory" banner so you can address it in your summary or cover letter.
// lib/analysis.ts -> analyzeExperience() (pure TypeScript)
function parseMonthDate(value: string): number | null {
if (/present|current/i.test(value)) return Date.now()
const iso = value.match(/(\d{4})-(\d{2})/)
if (iso) return new Date(Number(iso[1]), Number(iso[2]) - 1, 1).getTime()
const parsed = Date.parse(value) // "Mar 2021", "March 2021"...
return Number.isNaN(parsed) ? null : parsed // unparseable -> skip, never guess
}
// Roles arrive newest-first. Compare each role's start with the END of the
// role before it; a hole of 3+ months is flagged (first gap only).
for (let i = 0; i < roles.length - 1; i++) {
const earlierRoleEnd = parseMonthDate(roles[i + 1].endDate)
const laterRoleStart = parseMonthDate(roles[i].startDate)
if (earlierRoleEnd === null || laterRoleStart === null) continue
const gapMonths = Math.round((laterRoleStart - earlierRoleEnd) / (1000 * 60 * 60 * 24 * 30))
if (gapMonths >= 3) {
gap = { detected: true, /* duration, period, between */ }
break
}
}
// The gap is advisory. It is displayed, but never subtracted from a score.Keyword and knockout match
ATS keyword scans look for the terms an employer listed. This stage checks each requirement against everything in your resume.
- Runs on
- Your device (browser), using extracted data
- Method
- LLM-extracted requirements, matched by code
- Available to
- Everyone, including guests
APIs and libraries used in this step
openai/gpt-4o-mini
AI model · extracts requirements and knockouts in /api/extract
matchRequirements()
Our code · lib/analysis.ts, pure TypeScript
Algorithm
- Build one lowercase corpus from your headline, skills list, and every role title and bullet.
- For each requirement, test a direct match: does the corpus contain the full requirement text?
- Otherwise test a token match: split the requirement into words longer than two characters, drop stopwords, and check whether any remaining word appears.
- Report matched counts separately for must-haves and nice-to-haves. Knockout criteria from the job are listed as a notice. They are not checked against your resume.
Novel solution 6: two-tier matching with a shared corpus
// lib/analysis.ts -> matchRequirements() (pure TypeScript)
const STOPWORDS = new Set(['and', 'the', 'with', 'for', 'years', 'related', 'etc', 'plus'])
function normalizeTokens(value: string): string[] {
return value
.toLowerCase()
.replace(/[^a-z0-9+\s]/g, ' ') // keep "+" so c++ and "c#"-style terms survive
.split(/\s+/)
.filter((token) => token.length > 2)
}
export function matchRequirements(resume: ExtractedResume, requirements: ExtractedRequirement[]) {
// One searchable corpus: headline + skills + every role title and bullet.
const corpus = [
resume.headline,
resume.skills.join(' '),
...resume.roles.flatMap((role) => [role.title, ...role.bullets]),
].join(' ').toLowerCase()
return requirements.map((requirement) => {
const tokens = normalizeTokens(requirement.label).filter((t) => !STOPWORDS.has(t))
const directMatch = corpus.includes(requirement.label.toLowerCase())
const tokenMatch = tokens.length > 0 && tokens.some((token) => corpus.includes(token))
return { skill: requirement.label, matched: directMatch || tokenMatch }
})
}This matcher is intentionally lenient: "5+ years React" matches if the word "react" appears anywhere. It is a keyword scan, which is exactly what simple ATS filters do. Judging whether you truly have five years of React is the job of Stage 4.
Deep semantic match
Moves beyond keywords. A model reads the evidence in your resume and grades how well each requirement is genuinely met.
- Runs on
- ValidateCV server, then an AI model
- Method
- LLM judge at temperature 0, scored by code
- Available to
- Jobseeker and Career Coach plans
APIs and libraries used in this step
POST /api/assess
Server API · Next.js Route Handler
generateObject()
AI model · AI SDK, schema-validated output
Vercel AI Gateway
AI model · plain model ID
openai/gpt-4o-mini
AI model · the judge
weighted scoring loop
Our code · plain TypeScript, not the model
The 0 to 3 grading rubric
| Score | Meaning |
|---|---|
| 0 | No evidence anywhere in the resume |
| 1 | Weak or tangential evidence |
| 2 | Solid, directly relevant evidence |
| 3 | Evidence that meets or exceeds the requirement |
Algorithm
- Send the structured resume and every must-have and nice-to-have requirement (still redacted).
- The model returns, per requirement: a 0 to 3 score, a verbatim evidence quote from your resume, and a one-sentence reason. Quotes make each grade auditable.
- Code weights each score: must-haves count double. The maximum for each requirement is 3 times its weight.
- The overall score is the weighted total divided by the weighted maximum, as a percentage.
overall = round( Σ(score × weight) ÷ Σ(3 × weight) × 100 )
Worked example
| Requirement | Type | Score | Weight | Points | Max |
|---|---|---|---|---|---|
| React in production | Must-have | 3 | 2 | 6 | 6 |
| TypeScript | Must-have | 2 | 2 | 4 | 6 |
| GraphQL | Nice-to-have | 1 | 1 | 1 | 3 |
| Total | 11 | 15 |
11 ÷ 15 = 0.733, so the overall score is 73%.
Novel solution 7: the model grades, code calculates
Language models are unreliable at arithmetic and drift between runs. Constraining the model to a bounded integer plus evidence, then doing the maths in TypeScript, makes the formula auditable and repeatable.
// app/api/assess/route.ts (Stage 4: LLM judges, TypeScript does the maths)
const requirementJudgementSchema = z.object({
requirement: z.string(),
// 0 = no evidence, 1 = weak, 2 = solid match, 3 = meets or exceeds
score: z.number().int().min(0).max(3),
evidenceQuote: z.string(), // verbatim resume quote, or "" when score is 0
reasoning: z.string(), // one sentence
})
const { object } = await generateObject({
model: 'openai/gpt-4o-mini',
schema: z.object({ requirementJudgements: z.array(requirementJudgementSchema) }),
temperature: 0,
prompt: '<rubric + requirements + redacted resume>',
})
// The model NEVER computes the percentage. Must-haves weigh double.
const mustHaveLabels = new Set(jobDescription.mustHave.map((r) => r.label))
let weightedTotal = 0
let weightSum = 0
for (const judgement of object.requirementJudgements) {
const weight = mustHaveLabels.has(judgement.requirement) ? 2 : 1
weightedTotal += judgement.score * weight
weightSum += 3 * weight // 3 is the maximum score per requirement
}
const overallScore = weightSum > 0 ? Math.round((weightedTotal / weightSum) * 100) : 0Status. The /api/assess endpoint exists, but the report currently shows Stage 4 as a locked preview and does not call it yet. The AI rewrite suggestions and tone analysis described in that preview are planned and have no code behind them today.
Every API in one place
This is the complete list of what the simulator touches. There is no database, no file storage and no third-party data API in the scan path.
| API or library | Used for | Where it runs | What it sees |
|---|---|---|---|
| File API | Reading the uploaded file | Your device | Original file (never uploaded) |
| pdfjs-dist (legacy) | PDF to text | Your device | Original text |
| mammoth | DOCX to text | Your device | Original text |
| lib/redact-pii.ts | Removing personal details | Your device | Original text, outputs redacted text |
| sessionStorage | Passing redacted text to the report page | Your device | Redacted text only |
| POST /api/extract | Structured extraction | ValidateCV server | Redacted text |
| POST /api/assess | Stage 4 semantic judging | ValidateCV server | Redacted structured JSON |
| AI SDK generateObject + zod | Schema-validated model output | ValidateCV server | Redacted text |
| Vercel AI Gateway to openai/gpt-4o-mini | The language model | Model provider | Redacted text, subject to the provider's data policy |
Known limitations
We would rather you hear these from us than discover them. Each is a place where the current build falls short of what the interface might imply.
Redaction is heuristic.
It relies on section headings, patterns and word lists, not a trained entity recogniser. Unusual layouts, two-column PDFs and non-standard headings can hide structure. Check the Inspect your data panel after each scan.
Scanned PDFs cannot be read.
We read the PDF text layer. Image-only PDFs have none, and there is no OCR step.
The special-character check over-counts bullets.
The standard bullet character is in the decorative glyph list, so a resume with many bullet points can trigger a warning that is not a real problem.
Keyword matching is lenient and word-based.
Any single meaningful word from a requirement counts as a match, and matching is substring-based, so short terms can match inside longer words.
Years of experience and dates come from a model.
Extraction can be wrong on unusual date formats. Gap detection also assumes roles are listed newest first.
Stage 4 is not yet live in the report.
See the status note under Stage 4.
See it run on your own resume
Now that you know exactly what happens to your data and how each score is produced, run a free scan.
Run ATS Simulator