OpenAI’s GPT‑5.6 family includes three models: Sol, Terra and Luna. GPT‑5.6 introduces three model tiers, multiple reasoning-effort settings and two processing speeds. These options can make it difficult to predict which configuration will deliver the required result at the lowest total cost.
The most expensive model might not be the best and the “cheapest” model might end up being more expensive because of retries etc.
This article will help to explain the calculator above.
Limits of the Calculator:
One naming detail matters for developers: the API model name gpt-5.6 routes to Sol. It does not automatically select the cheapest model for each request. Terra and Luna must be selected explicitly with gpt-5.6-terra or gpt-5.6-luna.
These estimates exclude tool-call fees, cache writes and long-context surcharges.
The output price does not apply only to the answer users can see. OpenAI models may generate hidden reasoning tokens before or between visible outputs, and those reasoning tokens are billed at the model’s output-token rate.
When a GPT‑5.6 request contains more than 272,000 input tokens, OpenAI charges twice the normal input price and 1.5 times the normal output price for the entire request.
This matters for long coding sessions or conversation histories.
Cached reads receive a 90% discount. However, GPT‑5.6 cache writes cost 1.25 times the normal uncached input rate.
Caching is most valuable when the same large prompt prefix such as documentation will be reused.
Batch and Flex processing offer discounted rates for eligible workloads. See OpenAI’s API pricing for availability and current rates.
Which 5.6 Model and Effort to Use
In our own internal use we have found…
Luna - Luna at high or extra-high effort is a practical daily model but with checks needed for anything complicated.
Terra - Seems to be a reliable model that behaves as you would expect with Medium or High effort.
Sol - Especially when High or Ultra effort are selected can significantly overthink a prompt and solve the task in a more complex way than necessary.
Model selection is how “smart” the model is.
Effort is how “hard” the model will work.
Model choice and effort explained like hiring a plumber
The model represents the plumber’s level of expertise and the scale of work they can handle.
The effort setting controls how thoroughly they plan and investigate before starting and analyze after completing.
Effort determines how deeply the model examines the job:
A stronger model can handle a bigger job. Higher effort tells that model to think more carefully about how the job should be done and analyze their work after they completed it.
You would not pay a specialist to fix a loose faucet. But you also would not want someone guessing when water is leaking through your ceiling.
Use the cheapest model and effort level likely to do the job correctly. Choose a stronger model or higher effort when a wrong answer would cost more to fix.
Use Originality.ai’s free Sol vs Luna vs Terra calculator to compare estimated API costs using your own token volume, request frequency and caching assumptions.

Publishers and creators are responding to the use of AI crawlers by implementing standards to tell AI that it can’t train on their content. Are Noai and Noimageai tags being widely adopted? Find out and track adoption in our live dashboard and study.