Companies creating generative AI systems typically assess models using the same metric: cost per million tokens. The figure appears on all pricing pages, making it a common element in spreadsheets. However, production workloads do not acquire tokens.
They purchase results: a resolved assistance ticket, a finished research report, an accurate financial statement. On the path between the pricing page and the outcome, there are factors the sticker price overlooks: the accuracy of the model, the number of tokens required, and for goal-oriented tasks, the number of iterations needed, as each iteration resends the evolving conversation.
This post presents the results of an open-source benchmarking framework that evaluates these factors in OpenAI models on Amazon Bedrock (gpt-21100-luna, gpt-212026-terra, and gpt-215.6-sol) and two cost-effective models available on the OpenAI API (gpt-5.4-mini and gpt-5.4-nano). We selected the last two options as cost-effective starting points for numerous teams, rather than comparable generational counterparts, since “we currently operate mini or nano models.” The question we frequently come across is, “Is it worth purchasing a more recent model of Amazon Bedrock?” Our main emphasis is on addressing three inquiries:.
