Doubling a transformer's context window from 100k to 200k tokens does what to attention's compute cost, and why?