| |
Large language models can effectively compress responses into "telegraphese" (terse, abbreviated text mimicking 19th-century telegraph style) while maintaining or improving downstream accuracy, achieving 25-49% token savings across different model families. Testing on 1,300 benchmark questions shows that models reading compressed records perform nearly as well as or better than those reading plaintext (recovery ratios 0.99-1.10), suggesting this compression technique leverages language patterns already present in LLM training data. The technique is most valuable when the compressed output is consumed by another model rather than a human user.
Read Full Article →
← More Tech news