Prompt Caching & Cost

Key idea: You pay for every token in and out. Caching makes a repeated prompt start much cheaper, if you build prompts for it.

  1. Ask two questions in a row. The second request reads most of its prompt from the cache and costs less.
  2. Turn off Cache-friendly prompt and ask again. A timestamp at the top breaks the cache every time.
  3. Turn off the document. The prompt drops below ~1,024 tokens and can't be cached at all.
  4. Switch the answer length to Detailed. Output tokens cost about 4x input, so watch the Answer slice grow.
  5. Compare Small and Large models on the same question, and check the cost per 1,000 requests.
Try a prompt
Enter to send · Shift+Enter for a new line

Change how it works, then send again

usageSystem promptinstructionsDocumentcustomer agreementChat historyearlier turnsQuestionyour messagePromptall input tokensPrefix cacheseen this start?Modelreads, then writesgpt-4.1-nanoAnsweroutput tokensBilltokens × price
Send a message to watch it run

Cost

Ask a question to see where every token and dollar of the request goes.

Behind the scenes

Send a message and every step the system takes will show up here.