Description of the Corpus
We take the legal and ethical implications of using AI in Human Resources very seriously. The data we use inherently excludes any personal information. When context is provided, it is generated using generative Large Language Models, so all the information is produced without involving personal/company information or geographic location.
The specific details about the datasets, including file formats and examples, will be published here when the sample set is released.
Task 1: Contextualized Job-Person Matching
The dataset reuses and extends the 2026 data-generation pipeline: synthetic, multilingual job vacancies and résumés containing no personal information, covering English, Spanish and, new for this edition, French. Subtask 1.2 additionally provides evidence-unit annotations for a subset of job offer–candidate pairs.
Task 2: Paragraph-to-Skill Ranking under Negation
The dataset consists of short paragraphs of job-market text paired with ESCO skills annotated for polarity (required, negated, or neither). The training set is deliberately kept small, as the negation/polarity signal is scarce; participants may freely enrich it with other skill-extraction datasets.
Both datasets are annotated by expert annotators with professional experience in Human Resources.