No datasets match this combination of filters. Remove a filter from the query line above, or add the dataset you were looking for.
A living catalog of cultural-NLP benchmarks, coded by taxonomy branch :: category, representational mode, languages, regions, evaluated models, and evaluation-protocol flags.
No datasets match this combination of filters. Remove a filter from the query line above, or add the dataset you were looking for.
1 · The coding schema. Every dataset is coded on the schema of the accompanying meta-analysis: taxonomy branch :: category pairs (Ideational / Linguistic / Social × 14 cultural elements, after Liu et al.), the representational mode assigned with the translation-invariance test — if the item were faithfully translated into another language, the gold answer holds (C, culture-isolating), changes (CL, culture–language entangled), or becomes void (L, linguistic form only) — and the four evaluation-protocol flags: refusal, agentic, safety, robustness. Hover the i icons next to any filter, or any mode/branch/category option, for the exact coding criterion of that field.
2 · Annotations are never collapsed. A dataset carries every accepted annotation side by side: the
curated baseline from the released coding, the study's validation annotators (A1, A2, A3, and Claude), and community
contributions. Filters match a dataset if any of its annotations matches, and displayed values carry annotator counts —
e.g. Ideational :: Knowledge (3) means three annotators assigned that pair. Inter-annotator
disagreement stays visible in the interface rather than being adjudicated away; expand any row to compare annotations.
3 · Contributing. Two paths, both public GitHub issue forms mirroring the schema: Add a new dataset (top right) is for datasets not yet listed — an automatic check compares your submission against the catalog by dataset name, paper title, and paper link, and redirects likely duplicates to the annotation flow. + annotate, next to every dataset, adds your own independent annotation of that dataset, prefilled with its name and title. Each annotation is published under your chosen name or alias, with one annotation per alias per dataset — the automatic check verifies your alias is still free; to revise your earlier annotation, comment on your original submission instead. Besides the schema fields, your annotation can also confirm or correct the dataset's languages, cultural regions, and evaluated models; accepted values join the dataset's filterable metadata.
4 · Review and publication. Every submission notifies the maintainers by e-mail and is reviewed against the coding manual (usually within two weeks); reviewers may ask follow-up questions in the issue thread. Accepted entries are added to the released data, the catalog is rebuilt, and the issue is closed with a link to the entry. Declined submissions receive a reason. The full dataset — including all annotations — can be exported as CSV from any filtered view above.