About DAT Scoring
CAP's DAT scoring models calculate the average semantic distance between words in each DAT response using word embeddings.
For English data, we recommend using the Olson et al., 2021 model, which computes semantic distance using the GloVe 840B semantic space.
For data in any other language, we recommend using the Haase et al., 2025 model (S-DAT). All languages and scripts are supported, although results might vary across languages due availability of training data.
How do I use DAT Scoring?
To have your DAT responses automatically scored, upload a CSV that includes a column named 'response', which contains the participant's response. DAT responses should be formatted as one set of at least two unique, comma-separated words per cell. Responses with duplicate words cannot be scored. The algorithm is case sensitive for the 'response' column name, but it does not matter if other columns are present in the CSV.
The algorithm will add two new columns to your CSV: one named 'prediction' which reflects the model's score and a column named 'modelname', which is added to reflect the model providing the score.