eval suite
pr0xyh0rse evals are being developed to test what standard benchmarks often miss: continuity, state custody, boundary handling, relational consistency, nonhuman perception, and behaviour under generative pressure.
The goal is not to reward models for sounding right once. The goal is to inspect whether they can preserve the important things across context, pressure, ambiguity, and time.