All resources
Topic

model-evaluation-research

1 resources across 1 kinds

References

  1. Primary research from the UK AI Security Institute (DSIT): frontier-model cyber-capability evaluations, safeguard and jailbreak testing, and incident reports from their own agent testing.

    Open ↗