Description of the Corpus

Task 1: Contextualized Job-Person Matching

The dataset reuses and extends the 2026 data-generation pipeline: synthetic, multilingual job vacancies and résumés containing no personal information, covering English, Spanish and, new for this edition, French. Subtask 1.2 additionally provides evidence-unit annotations for a subset of job offer–candidate pairs.

Task 2: Paragraph-to-Skill Ranking under Negation

The dataset consists of short paragraphs of job-market text paired with ESCO skills annotated for polarity (required, negated, or neither). The training set is deliberately kept small, as the negation/polarity signal is scarce; participants may freely enrich it with other skill-extraction datasets.

Both datasets are annotated by expert annotators with professional experience in Human Resources.