[ MANIFEST · #EVALUATION ]
DATA[NEW]
What context, translation specialists, and JSON schemas changed
A paired evaluation of translation context, specialist and general models, structure-aware parsing, and schema-constrained decoding on reasoning data.
When the translator starts solving the problem
An evaluation of task execution failures, chunking, and translation quality when translating long reasoning traces.