ValidateCV logoValidateCV
Transparency report

How the ATS simulator actually works

The rubric behind every score, the step-by-step algorithm for each of the four assessment stages, which API is called at each point, and the real code for the parts we built ourselves.

Read this first. ValidateCV is a simulator, not a copy of any specific employer's ATS. Vendors such as Workday, Greenhouse and Lever do not publish their parsing rules. Our rubric encodes widely documented ATS failure modes. We mark every heuristic as a heuristic and list what is not finished yet in Known limitations.

The pipeline at a glance

One scan is a single run through the steps below. Green panels run entirely in your browser. The blue panel is the only place text leaves your device, and by then it has been redacted. The dashed panel is the paid stage.

Your device

Browser only. Nothing here touches the network.

  1. Add resume and job description

    Drag-and-drop PDF or DOCX, plus pasted text or an uploaded file for the job.

    File API

  2. Extract text from the files

    PDF text layer via a Web Worker; DOCX raw text. 20-second timeout guard.

    pdfjs-dist (legacy) / mammoth

  3. Redact personal details

    Emails, addresses, phone numbers and the name line are replaced with tokens.

    lib/redact-pii.ts (regex)

  4. Hold the redacted text for this tab

    Saved so the report page can read it. Cleared when the tab closes.

    sessionStorage

ValidateCV server and AI model

Receives redacted text only. Stores nothing.

  1. Receive redacted text

    Validates the request body with a schema, then forwards it for extraction.

    POST /api/extract (Next.js Route Handler)

  2. Turn text into structured data

    Two parallel model calls return typed JSON: roles, skills and dates from the resume; must-haves and knockouts from the job.

    AI SDK generateObject + zod, via Vercel AI Gateway (openai/gpt-4o-mini, temperature 0)

Your device again: free report

Structured JSON comes back. Scoring runs locally, with no more AI calls.

  1. Stage 1: Parsing and readability

    Five formatting checks on the redacted text, averaged into a score.

    analyzeParsing()

  2. Stage 2: Experience and recency

    Years, current role and employment-gap detection from the extracted roles.

    analyzeExperience()

  3. Stage 3: Keyword and knockout match

    Each job requirement is searched for in the resume corpus.

    matchRequirements()

Stage 4: Deep semantic match

Endpoint built and tested. Shown as a locked preview in the report today.

  1. Judge each requirement 0 to 3

    The model scores and quotes evidence. A TypeScript function does the weighted maths.

    POST /api/assess, generateObject via AI Gateway (temperature 0)

The ATS rubric

The report has four stages. Stages 1 to 3 are the free Guest Check: they flag red flags and never call an AI model to score you. Stage 4 is the only stage where a model forms a judgement.

The four ATS assessment stages and how each is scored
StageWhat it checksMethodHow it is scoredTier
1. Parsing and readabilityWhether a parser can read your layout: length, columns, headers, tables, special characters.Deterministic codeFive checks, each pass (100), warning (60) or fail (20). Score is the mean.Free
2. Experience and recencyTotal years, most recent role, and unexplained employment gaps.Model extracts dates; code analyses themInformational. Gaps are advisory and never lower a score.Free
3. Keyword and knockout matchWhether each must-have and nice-to-have requirement appears in your resume. Lists knockout criteria.Model extracts requirements; code matches themMatched count per group. Knockouts are shown, not scored.Free
4. Deep semantic matchWhether your experience genuinely satisfies each requirement, beyond exact words.Model judges; code calculatesEach requirement scored 0 to 3, weighted, converted to a 0 to 100 score.Paid

Design principles behind the rubric

  • Models extract and judge; code scores. No percentage you see is invented by a language model. Models produce structured facts or a 0 to 3 grade, and plain TypeScript turns those into numbers.
  • Temperature 0 everywhere. Extraction and judging are classification tasks, so the same input should give the same output.
  • Red flags, not verdicts. Stages 1 to 3 tell you what a bot could trip on. They do not predict whether you will be hired.
Before the stages

Intake: parsing and PII redaction

Your files become plain text on your device, and personal details are stripped before anything is stored or sent.

Runs on
Your device (browser)
Method
Libraries and regular expressions
Available to
Everyone, including guests

APIs and libraries used in this step

  • File.arrayBuffer()

    Browser API · reads the upload; the file is never sent anywhere

  • pdfjs-dist (legacy build)

    Library (in browser) · PDF text layer, runs in a Web Worker

  • mammoth

    Library (in browser) · DOCX raw text

  • lib/redact-pii.ts

    Our code · section-aware, one-way redaction

  • sessionStorage

    Browser API · redacted text only, per tab

Algorithm

  1. Detect the file type from the MIME type or the extension, then route to the PDF or DOCX parser.
  2. Extract text locally. For PDFs, group text fragments into real lines by their coordinates, then drop page numbers and any top or bottom line that repeats across pages (running headers and footers). If no text exists (a scanned image), the run continues with a placeholder and you are warned.
  3. Race the parser against a 20-second timer. A stalled PDF worker can never freeze the page.
  4. Run redactResume() on the resume and the lighter redactPII() on the job description.
  5. Write only the redacted text to sessionStorage, then navigate to the report.

Novel solution 1: a parser that cannot hang

Chrome stalled on the current PDF.js build because its worker relies on JavaScript APIs Chrome does not ship yet. We import the legacy build, which bundles polyfills, and we wrap every parse in a timeout race.

lib/parse-file.ts
// lib/parse-file.ts  (runs in the browser)
const PARSE_TIMEOUT_MS = 20_000

export async function parsePdfToText(file: File): Promise<string> {
  // The legacy build ships polyfills Chrome may lack. The modern build
  // crashes its worker there and getDocument() never settles.
  const pdfjsLib = await import('pdfjs-dist/legacy/build/pdf.mjs')
  pdfjsLib.GlobalWorkerOptions.workerSrc = new URL(
    'pdfjs-dist/legacy/build/pdf.worker.mjs',
    import.meta.url,
  ).toString()

  const doc = await pdfjsLib.getDocument({ data: await file.arrayBuffer() }).promise
  const pageTexts: string[] = []
  for (let n = 1; n <= doc.numPages; n++) {
    const content = await (await doc.getPage(n)).getTextContent()
    pageTexts.push(content.items.map((i) => ('str' in i ? i.str : '')).join(' '))
  }
  return pageTexts.join('\n\n').replace(/[ \t]+/g, ' ').trim()
}

export async function parseResumeFile(file: File): Promise<string> {
  const isPdf = file.type === 'application/pdf' || file.name.toLowerCase().endsWith('.pdf')
  const parsing = isPdf ? parsePdfToText(file) : parseDocxToText(file)

  // A worker that fails to start can leave parsing pending forever.
  // Race it against a timer so the UI can never freeze.
  let timeoutId: ReturnType<typeof setTimeout>
  const timeout = new Promise<never>((_, reject) => {
    timeoutId = setTimeout(() => reject(new Error('Reading your file took too long.')), PARSE_TIMEOUT_MS)
  })
  try {
    return await Promise.race([parsing, timeout])
  } finally {
    clearTimeout(timeoutId!)
  }
}

Novel solution 2: structure-aware, one-way redaction

Redaction happens before any network request, so the server and the model never receive the original identifiers. The resume is split into sections first, because what counts as personal depends on where it sits. The rules run in this order:

  1. Everything before the first section heading (name, title, subtitle, contact block) collapses into one header token.
  2. References, referees and personal-details or declaration sections are removed whole.
  3. In work experience, only the job title, the dates and the task text survive. Employer names, locations and contacts are redacted, and when a line cannot be positively identified as a title or a date it is treated as an employer.
  4. Employer names (including short forms such as "Acme") and the candidate's name are then scrubbed from the rest of the document, so they cannot reappear in a task bullet or footer.
  5. Emails, phone numbers, links and handles, street addresses, suburb and state, and labelled personal details (date of birth, gender, marital status, nationality, visa, licence, TFN) are redacted.
  6. Leftover header and footer lines such as "Name | Page 2" are removed.

Redaction is irreversible from the output alone. Matches record a type and nothing else, every placeholder is a fixed string with no numbering, hash, length or partial digits, a removed block becomes a single token so its size does not leak, and the original text is never written to storage.

lib/redact-pii.ts (excerpt)
// lib/redact-pii.ts  (runs in the browser, BEFORE any network call)  - excerpt
export function redactResume(text: string): RedactResumeResult {
  const lines = text.split('\n').map((line) => line.replace(/\s+/g, ' ').trim())
  const headings = lines.map(detectHeading) // "Experience", "References", ...

  // 1. Everything before the first section heading is the header block.
  const firstSection = headings.findIndex((kind) => kind !== null)
  const preSection = lines.slice(0, firstSection)
  const nameTerms = namesFrom(preSection)
  const out: string[] = []
  if (preSection.some(Boolean)) out.push('[REDACTED_HEADER]') // one token: size does not leak

  for (let i = firstSection; i < lines.length; i++) {
    const heading = headings[i]
    if (heading === 'references' || heading === 'personal') {
      removedBlock = heading            // 2. drop the whole section
      out.push('[REDACTED_' + heading.toUpperCase() + ']')
      continue
    }
    if (removedBlock) continue

    if (section === 'experience' && entry.mode === 'header') {
      // 3. job header lines keep title + dates only
      const result = redactJobHeaderLine(lines[i], matches)
      companyTerms.push(...result.companyNames.flatMap(companyVariants))
      out.push(result.line)
      continue
    }
    out.push(redactEntities(lines[i], matches)) // email, phone, links, address...
  }

  // 4. Names and employers can reappear in task bullets: scrub them everywhere.
  let scrubbed = scrubGlobally(out, nameTerms, 'name', matches)
  scrubbed = scrubGlobally(scrubbed, companyTerms, 'company', matches)
  // 5. Remove leftover header/footer lines such as "Name | Page 2".
  return finish(scrubbed.map((l) => (isFooterLine(l) ? '[REDACTED_FOOTER]' : l)))
}

// "Acme Technologies Pty Ltd" is also written "Acme Technologies" and "Acme".
function companyVariants(name: string): string[] {
  const withoutLegal = stripLegalSuffix(name)
  const core = stripGenericCompanyWords(withoutLegal)
  return [name, withoutLegal, core]
}

// Unsure whether a header segment is a job title? Treat it as an employer.
function classifySegment(text: string): 'title' | 'company' {
  const isTitle = !LEGAL_SUFFIX.test(text) && TITLE_WORDS.test(text) && wordCount(text) <= 6
  return isTitle ? 'title' : 'company'
}

Honest scope of the privacy claim. Redaction is a set of heuristics, not a trained entity recogniser. It does not yet remove university names or the names of people mentioned inside task text, and a resume with no recognisable section headings only loses its leading contact lines and labelled personal details. The redacted text is sent to our server and on to the AI model provider through Vercel AI Gateway. We have no database and our routes do not write your text anywhere, but the provider's own data policy applies to what it receives.

Before the stages

Extraction: unstructured text to structured data

Stages 2 and 3 need facts, such as dates, roles and requirements, not raw paragraphs. One server call turns both documents into typed JSON.

Runs on
ValidateCV server, then an AI model
Method
LLM extraction at temperature 0
Available to
Everyone, including guests

APIs and libraries used in this step

  • POST /api/extract

    Server API · Next.js Route Handler

  • generateObject()

    AI model · AI SDK, schema-validated output

  • zod

    AI model · defines and enforces the output schema

  • Vercel AI Gateway

    AI model · plain model ID, no provider SDK

  • openai/gpt-4o-mini

    AI model · the model used

Algorithm

  1. The report page reads the redacted text from sessionStorage and POSTs both documents to /api/extract.
  2. The route validates the body with zod and returns a 400 for anything malformed.
  3. Two independent generateObject calls run in parallel with Promise.all: one for the resume, one for the job description.
  4. Each response is forced to match its zod schema, so the client always receives well-formed JSON rather than free text.
  5. The JSON is returned to the browser. Nothing is persisted.

Novel solution 3: schemas as the contract

The model is never asked for a score. It fills a schema, and the field descriptions double as instructions. The job description schema separates must-haves from nice-to-haves and isolates knockouts: hard filters such as work rights or clearance that an ATS auto-rejects on. Prompts are abbreviated below.

app/api/extract/route.ts
// app/api/extract/route.ts  (ValidateCV server -> Vercel AI Gateway)
const resumeSchema = z.object({
  headline: z.string(),
  roles: z.array(z.object({
    title: z.string(),
    employer: z.string(),
    startDate: z.string(),   // "2021-03" or "Mar 2021"
    endDate: z.string(),     // "2023-08" or "Present"
    bullets: z.array(z.string()),
  })),
  skills: z.array(z.string()),
  totalYearsExperience: z.number(),
})

const jobDescriptionSchema = z.object({
  title: z.string(),
  mustHave: z.array(requirementSchema),
  niceToHave: z.array(requirementSchema),
  knockouts: z.array(z.string()),   // hard filters: visa, clearance, mandatory degree
  keywords: z.array(z.string()),
})

// The two extractions are independent, so they run in parallel.
const [jdResult, resumeResult] = await Promise.all([
  generateObject({
    model: 'openai/gpt-4o-mini',   // plain Gateway model ID, no provider SDK
    schema: jobDescriptionSchema,
    temperature: 0,                // parsing task: same input -> same output
    prompt: '<extraction instructions + redacted job description>',
  }),
  generateObject({
    model: 'openai/gpt-4o-mini',
    schema: resumeSchema,
    temperature: 0,
    prompt: '<extraction instructions + redacted resume text>',
  }),
])
app/analysis/page.tsx
// app/analysis/page.tsx  (the only network call the free report makes)
const response = await fetch('/api/extract', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({ resumeText, jobDescriptionText }),  // both already redacted
})
const data = await response.json()   // { resume, jobDescription }  structured JSON

// Everything after this line (Stages 1-3) is computed locally in the browser.
Stage 1

Parsing and readability

Can a machine read your resume the way you intended? This stage inspects the extracted text for the layout problems that scramble real parsers.

Runs on
Your device (browser)
Method
Deterministic code, no AI
Available to
Everyone, including guests

APIs and libraries used in this step

  • analyzeParsing()

    Our code · lib/analysis.ts, pure TypeScript

The five checks

Stage 1 checks and their thresholds
CheckPassWarningFail
Resume length400 to 800 words250 to 399 or 801 to 1,000Anything else
Single-column layout15% or fewer of lines look interleavedMore than 15% of lines have 4+ space gaps or exceed 220 charactersNever fails (it is a proxy)
Standard section headers3 or 4 of Summary, Experience, Education, Skills2 recognised0 or 1 recognised
Tables and text boxesUp to 5 tab runs or pipe separatorsMore than 5Never fails
Fonts and special characters20 or fewer decorative glyphsMore than 20 bullets, stars, hearts or emojiNever fails

Scoring algorithm

  1. Run the five checks on the redacted resume text.
  2. Map each status to points: pass 100, warning 60, fail 20.
  3. Average the five values and round. Below 70 shows a critical "likely unreadable" flag. 70 to 87 shows a minor-risk notice. 88 and above shows no banner.

Novel solution 4: detecting columns from their symptoms

We receive text, not a rendered page, so we cannot see columns. We detect what columns do to extracted text instead: two columns read line by line produce very long lines and wide internal whitespace. Likewise, fonts are invisible in extracted text, so the font check looks for glyphs known to garble.

lib/analysis.ts (analyzeParsing)
// lib/analysis.ts  ->  analyzeParsing()   (pure TypeScript, no network)
const STATUS_SCORE = { pass: 100, warning: 60, fail: 20 }

// Multi-column PDFs extract as interleaved text: long lines or big
// whitespace gaps. We can't see the layout, so we detect its symptoms.
const lines = resumeText.split('\n').filter((line) => line.trim().length > 0)
const suspiciousLines = lines.filter((line) => /\s{4,}/.test(line) || line.length > 220).length
const columnStatus =
  lines.length === 0 ? 'warning'
  : suspiciousLines / lines.length > 0.15 ? 'warning'
  : 'pass'

// Section headers are matched by intent, not exact wording.
const expectedHeaders = [
  { label: 'Summary',    pattern: /\b(summary|profile|objective)\b/i },
  { label: 'Experience', pattern: /\b(experience|employment|work history)\b/i },
  { label: 'Education',  pattern: /\beducation\b/i },
  { label: 'Skills',     pattern: /\bskills\b/i },
]
const found = expectedHeaders.filter(({ pattern }) => pattern.test(resumeText))
const headerStatus = found.length >= 3 ? 'pass' : found.length === 2 ? 'warning' : 'fail'

// Final score: unweighted mean of the five check scores.
const score = Math.round(
  checklist.reduce((sum, item) => sum + STATUS_SCORE[item.status], 0) / checklist.length,
)
Stage 2

Experience and recency

How much experience you have, what your current role is, and whether there is a timeline gap a recruiter or filter would notice.

Runs on
Your device (browser), using extracted data
Method
LLM-extracted dates, analysed by code
Available to
Everyone, including guests

APIs and libraries used in this step

  • openai/gpt-4o-mini

    AI model · extracts roles and dates in /api/extract

  • analyzeExperience()

    Our code · lib/analysis.ts, pure TypeScript

Algorithm

  1. Years: taken from the model's totalYearsExperience estimate. This number is the model's, not calculated by us.
  2. Current role: the first role whose end date reads "present" or "current", otherwise the first role listed.
  3. Gap detection: parse each date to a timestamp, compare each role's start with the end of the role before it, and flag the first hole of three months or more.
  4. Unparseable dates are skipped rather than guessed.

Novel solution 5: advisory gaps

Gaps are real, common and often legitimate, so this check is deliberately incapable of lowering your score. It surfaces an "Advisory" banner so you can address it in your summary or cover letter.

lib/analysis.ts (analyzeExperience)
// lib/analysis.ts  ->  analyzeExperience()   (pure TypeScript)
function parseMonthDate(value: string): number | null {
  if (/present|current/i.test(value)) return Date.now()
  const iso = value.match(/(\d{4})-(\d{2})/)
  if (iso) return new Date(Number(iso[1]), Number(iso[2]) - 1, 1).getTime()
  const parsed = Date.parse(value)               // "Mar 2021", "March 2021"...
  return Number.isNaN(parsed) ? null : parsed    // unparseable -> skip, never guess
}

// Roles arrive newest-first. Compare each role's start with the END of the
// role before it; a hole of 3+ months is flagged (first gap only).
for (let i = 0; i < roles.length - 1; i++) {
  const earlierRoleEnd = parseMonthDate(roles[i + 1].endDate)
  const laterRoleStart = parseMonthDate(roles[i].startDate)
  if (earlierRoleEnd === null || laterRoleStart === null) continue

  const gapMonths = Math.round((laterRoleStart - earlierRoleEnd) / (1000 * 60 * 60 * 24 * 30))
  if (gapMonths >= 3) {
    gap = { detected: true, /* duration, period, between */ }
    break
  }
}
// The gap is advisory. It is displayed, but never subtracted from a score.
Stage 3

Keyword and knockout match

ATS keyword scans look for the terms an employer listed. This stage checks each requirement against everything in your resume.

Runs on
Your device (browser), using extracted data
Method
LLM-extracted requirements, matched by code
Available to
Everyone, including guests

APIs and libraries used in this step

  • openai/gpt-4o-mini

    AI model · extracts requirements and knockouts in /api/extract

  • matchRequirements()

    Our code · lib/analysis.ts, pure TypeScript

Algorithm

  1. Build one lowercase corpus from your headline, skills list, and every role title and bullet.
  2. For each requirement, test a direct match: does the corpus contain the full requirement text?
  3. Otherwise test a token match: split the requirement into words longer than two characters, drop stopwords, and check whether any remaining word appears.
  4. Report matched counts separately for must-haves and nice-to-haves. Knockout criteria from the job are listed as a notice. They are not checked against your resume.

Novel solution 6: two-tier matching with a shared corpus

lib/analysis.ts (matchRequirements)
// lib/analysis.ts  ->  matchRequirements()   (pure TypeScript)
const STOPWORDS = new Set(['and', 'the', 'with', 'for', 'years', 'related', 'etc', 'plus'])

function normalizeTokens(value: string): string[] {
  return value
    .toLowerCase()
    .replace(/[^a-z0-9+\s]/g, ' ')   // keep "+" so c++ and "c#"-style terms survive
    .split(/\s+/)
    .filter((token) => token.length > 2)
}

export function matchRequirements(resume: ExtractedResume, requirements: ExtractedRequirement[]) {
  // One searchable corpus: headline + skills + every role title and bullet.
  const corpus = [
    resume.headline,
    resume.skills.join(' '),
    ...resume.roles.flatMap((role) => [role.title, ...role.bullets]),
  ].join(' ').toLowerCase()

  return requirements.map((requirement) => {
    const tokens = normalizeTokens(requirement.label).filter((t) => !STOPWORDS.has(t))

    const directMatch = corpus.includes(requirement.label.toLowerCase())
    const tokenMatch = tokens.length > 0 && tokens.some((token) => corpus.includes(token))

    return { skill: requirement.label, matched: directMatch || tokenMatch }
  })
}

This matcher is intentionally lenient: "5+ years React" matches if the word "react" appears anywhere. It is a keyword scan, which is exactly what simple ATS filters do. Judging whether you truly have five years of React is the job of Stage 4.

Stage 4 (paid)

Deep semantic match

Moves beyond keywords. A model reads the evidence in your resume and grades how well each requirement is genuinely met.

Runs on
ValidateCV server, then an AI model
Method
LLM judge at temperature 0, scored by code
Available to
Jobseeker and Career Coach plans

APIs and libraries used in this step

  • POST /api/assess

    Server API · Next.js Route Handler

  • generateObject()

    AI model · AI SDK, schema-validated output

  • Vercel AI Gateway

    AI model · plain model ID

  • openai/gpt-4o-mini

    AI model · the judge

  • weighted scoring loop

    Our code · plain TypeScript, not the model

The 0 to 3 grading rubric

Stage 4 per-requirement score meanings
ScoreMeaning
0No evidence anywhere in the resume
1Weak or tangential evidence
2Solid, directly relevant evidence
3Evidence that meets or exceeds the requirement

Algorithm

  1. Send the structured resume and every must-have and nice-to-have requirement (still redacted).
  2. The model returns, per requirement: a 0 to 3 score, a verbatim evidence quote from your resume, and a one-sentence reason. Quotes make each grade auditable.
  3. Code weights each score: must-haves count double. The maximum for each requirement is 3 times its weight.
  4. The overall score is the weighted total divided by the weighted maximum, as a percentage.

overall = round( Σ(score × weight) ÷ Σ(3 × weight) × 100 )

Worked example

Worked Stage 4 example with three requirements
RequirementTypeScoreWeightPointsMax
React in productionMust-have3266
TypeScriptMust-have2246
GraphQLNice-to-have1113
Total1115

11 ÷ 15 = 0.733, so the overall score is 73%.

Novel solution 7: the model grades, code calculates

Language models are unreliable at arithmetic and drift between runs. Constraining the model to a bounded integer plus evidence, then doing the maths in TypeScript, makes the formula auditable and repeatable.

app/api/assess/route.ts
// app/api/assess/route.ts  (Stage 4: LLM judges, TypeScript does the maths)
const requirementJudgementSchema = z.object({
  requirement: z.string(),
  // 0 = no evidence, 1 = weak, 2 = solid match, 3 = meets or exceeds
  score: z.number().int().min(0).max(3),
  evidenceQuote: z.string(),   // verbatim resume quote, or "" when score is 0
  reasoning: z.string(),       // one sentence
})

const { object } = await generateObject({
  model: 'openai/gpt-4o-mini',
  schema: z.object({ requirementJudgements: z.array(requirementJudgementSchema) }),
  temperature: 0,
  prompt: '<rubric + requirements + redacted resume>',
})

// The model NEVER computes the percentage. Must-haves weigh double.
const mustHaveLabels = new Set(jobDescription.mustHave.map((r) => r.label))
let weightedTotal = 0
let weightSum = 0
for (const judgement of object.requirementJudgements) {
  const weight = mustHaveLabels.has(judgement.requirement) ? 2 : 1
  weightedTotal += judgement.score * weight
  weightSum += 3 * weight          // 3 is the maximum score per requirement
}
const overallScore = weightSum > 0 ? Math.round((weightedTotal / weightSum) * 100) : 0

Status. The /api/assess endpoint exists, but the report currently shows Stage 4 as a locked preview and does not call it yet. The AI rewrite suggestions and tone analysis described in that preview are planned and have no code behind them today.

Every API in one place

This is the complete list of what the simulator touches. There is no database, no file storage and no third-party data API in the scan path.

Complete inventory of APIs and libraries used by the ATS simulator
API or libraryUsed forWhere it runsWhat it sees
File APIReading the uploaded fileYour deviceOriginal file (never uploaded)
pdfjs-dist (legacy)PDF to textYour deviceOriginal text
mammothDOCX to textYour deviceOriginal text
lib/redact-pii.tsRemoving personal detailsYour deviceOriginal text, outputs redacted text
sessionStoragePassing redacted text to the report pageYour deviceRedacted text only
POST /api/extractStructured extractionValidateCV serverRedacted text
POST /api/assessStage 4 semantic judgingValidateCV serverRedacted structured JSON
AI SDK generateObject + zodSchema-validated model outputValidateCV serverRedacted text
Vercel AI Gateway to openai/gpt-4o-miniThe language modelModel providerRedacted text, subject to the provider's data policy

Known limitations

We would rather you hear these from us than discover them. Each is a place where the current build falls short of what the interface might imply.

  • Redaction is heuristic.

    It relies on section headings, patterns and word lists, not a trained entity recogniser. Unusual layouts, two-column PDFs and non-standard headings can hide structure. Check the Inspect your data panel after each scan.

  • Scanned PDFs cannot be read.

    We read the PDF text layer. Image-only PDFs have none, and there is no OCR step.

  • The special-character check over-counts bullets.

    The standard bullet character is in the decorative glyph list, so a resume with many bullet points can trigger a warning that is not a real problem.

  • Keyword matching is lenient and word-based.

    Any single meaningful word from a requirement counts as a match, and matching is substring-based, so short terms can match inside longer words.

  • Years of experience and dates come from a model.

    Extraction can be wrong on unusual date formats. Gap detection also assumes roles are listed newest first.

  • Stage 4 is not yet live in the report.

    See the status note under Stage 4.

See it run on your own resume

Now that you know exactly what happens to your data and how each score is produced, run a free scan.

Run ATS Simulator