Amazon's mid-tier multimodal model, reachable only through Bedrock, with a 300K context, a 5K output cap and a three-level reasoning effort switch.
Why this oneWhat it is actually for
Reach for Nova Pro when the application already lives in AWS and the conversation is about the bill rather than the benchmark. It is reachable only through Amazon Bedrock, so IAM roles, VPC endpoints, Guardrails, Knowledge Bases and Agents all work without a new vendor contract or a second egress path. A 300K context swallows whole document sets, screenshots and short video without chunking, and at $0.80 per million input tokens it sits well under the frontier models available behind the same endpoint. For document extraction, classification, summarisation and tool loops over well-defined tools, the capability gap rarely shows. For hard multi-step reasoning it does, and Bedrock makes swapping the model ID the cheapest fix.
What it isArchitecture, lineage, training
Amazon discloses almost nothing about how Nova Pro is built. The Bedrock model card and the Amazon Nova user guide give operating numbers - context window, output cap, modalities, cutoff - and stop there. Parameter count, dense versus mixture-of-experts, layer counts, tokenizer, training data and training compute are all Not disclosed. There is no paper and no weights release. Anything you read elsewhere about Nova Pro's parameter count is not coming from Amazon.
What Amazon does document is the family shape. Nova Version 1 has four understanding models - Micro (text only, 128K), Lite (300K), Pro (300K) and Premier (1M) - plus Canvas and Reel for image and video generation and Sonic for speech. The distillation relationships are stated explicitly: Nova Pro is a student of Premier and a teacher to Lite and Micro. That is the one real architectural hint on the page, and it implies a single family trained together with capability transferred downward rather than four unrelated models.
Nova Pro launched on 5 December 2024 with an October 2024 knowledge cutoff. It takes text, images and video in and emits text only; audio input is marked unsupported on its model card. Documents are handled natively in PDF, CSV, DOC, DOCX, XLS, XLSX, HTML, TXT and MD. Amazon claims support for over 200 languages with optimisation for 15 of them.
One thing to know before you start: Amazon has since shipped a Nova 2 line. The Nova 2 developer guide describes Nova 2 Lite with extended thinking and built-in web grounding and code interpreter tools, and the AWS Price List carries live SKUs for a Nova 2.0 Pro at roughly 1.7x Nova Pro's input rate and 3.4x its output rate. This page is about amazon.nova-pro-v1:0, the cheaper Version 1 model.
At a glanceSee it
How a Nova Pro request is shaped, cached, reasoned over and priced inside Amazon Bedrock.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| Context window | 300K tokens | Bedrock model card, Nova Pro |
| Max output tokens | 5K. The Nova request schema states "The maximum new tokens value allowed is 5K". The Amazon Nova user guide specification table says 10k for the same model - the two AWS pages disagree. | Bedrock model card; Nova complete request schema; Nova user guide |
| Input modalities | Text, image, video. Audio marked not supported. | Bedrock model card |
| Output modalities | Text only | Bedrock model card |
| Document formats | PDF, CSV, DOC, DOCX, XLS, XLSX, HTML, TXT, MD | Nova user guide, understanding model specifications |
| Knowledge cutoff | Oct 2024 | Bedrock model card |
| Launch date | Dec 05, 2024 | Bedrock model card |
| Languages | 200+, optimised for 15 | Nova user guide |
| Weights available | No. Bedrock-only API access, no licence offered. | Nova user guide, "you must access the models through an API using Amazon Bedrock" |
| Architecture and parameter count | Not disclosed | No AWS page states either |
| Prompt caching | Yes. Min 1K tokens per checkpoint, max 4 checkpoints per request, 5 minute TTL, 20K tokens maximum cached. | Bedrock model card, prompt caching table |
| Fine-tuning | Yes, on Bedrock. Training $0.008 per 1K tokens, custom model storage $1.95 per model per month. | Nova user guide; AWS Price List, AmazonBedrock us-east-1 |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
messages | Conversation turns. First turn must be user. Content blocks are text, image or video. | Required | An assistant turn at the end prefills the reply - useful for forcing a format opener. |
system | System prompt, an array of text blocks. | Optional | The main place to suppress preamble and pin output shape. Also a valid cache checkpoint location. |
inferenceConfig.maxTokens | Output cap. | Greater than 0, up to 5K. Default is dynamic. | Set it low and you truncate mid-sentence; the model does not compress to fit. |
inferenceConfig.temperature | Randomness. | 0.00001 to 1 inclusive. Default 0.7. | Capped at 1, unlike most APIs. There is no 0.0 - the floor is 0.00001. |
inferenceConfig.topP | Nucleus sampling. | 0 to 1 inclusive. Default 0.9. | AWS says alter either temperature or topP, not both. |
inferenceConfig.topK | Top-k truncation. | Docs conflict: the schema comment says "0 or greater, default 50", the prose says "between 0 and 128" and "the default value is that this parameter is not used". | Cuts the long tail. Because of the conflict, set it explicitly rather than relying on a default. |
inferenceConfig.stopSequences | Strings that halt generation. | Array of strings | Output is returned up to the match. Cheapest way to end a structured block. |
inferenceConfig.reasoningConfig.type | Turns internal reasoning on. | enabled or disabled. Default disabled. Nova Pro and Nova Lite only. | Off by default - if you assumed Nova Pro was thinking, it was not. |
inferenceConfig.reasoningConfig.maxReasoningEffort | How much reasoning to spend. | low, medium or high | At low and medium reasoning streams token by token; at high all reasoning arrives in one final chunk. |
toolConfig.tools | Tool specifications with JSON Schema input. | Tool name max 64 characters | Same schema as the Bedrock Converse tool-use API, so tools port across Bedrock models. |
toolConfig.toolChoice | Forces tool behaviour. | auto, any, or tool with a name | tool with a name is the only reliable way to get schema-shaped output from Nova. |
additionalModelRequestFields | Converse API escape hatch. | Object | topK and reasoningConfig must be nested here when using Converse. Put them at the top level and they do not take effect. |
| Prompt cache checkpoint | Marks a reusable prefix. | Up to 4 checkpoints, min 1K tokens each, 20K token ceiling, 5 minute TTL, accepted in system and messages | Cache reads bill at $0.20 per million against $0.80 - a 75% cut on the cached span. The Nova request schema page does not name the JSON field; the Bedrock prompt caching page does. |
response_format | Not supported. No JSON-mode or schema field appears in the Nova request schema. | Not supported | Use toolConfig.toolChoice set to a named tool and read the tool arguments. |
seed, logprobs, n, frequency and presence penalties | Not supported. None appear in the Nova request schema. | Not supported | For determinism, temperature 0.00001 plus a fixed prompt is the closest you get. For repetition, use stopSequences and prompt instructions. |
Three knobs carry the weight here. reasoningConfig is the big one because it is off by default and flipping it to enabled with medium effort changes what class of problem the model can finish. The cache checkpoint is second: any workload with a stable system prompt or a repeated document prefix gets a 75% discount on that span for free, since cache writes cost nothing. Third is additionalModelRequestFields - the single most common Nova bug is setting topK or reasoningConfig at the top level of a Converse call and quietly getting neither.
SamplingShaping the output distribution
Nova Pro's sampling surface is small and its ranges are unusual. Temperature is bounded at 1, not 2, and its floor is 0.00001 rather than 0 - there is no true greedy mode exposed. The default is 0.7, which is warmer than most enterprise APIs ship, so an out-of-the-box Nova Pro is noticeably chattier and more variable than the same prompt against a model defaulting to 0.2 or 0.3. Nucleus sampling defaults to 0.9. AWS is explicit that you should alter either temperature or topP but not both, which is a real recommendation and not boilerplate: with both moved you cannot attribute a behaviour change to either. topK is documented inconsistently - one part of the schema says the default is 50, another says it is unused - so set it yourself if you care. A sensible starting point for extraction and classification is temperature 0.00001 with topP and topK left alone; for drafting, leave the 0.7 default and adjust topP only if the output wanders.
ReasoningThinking, effort and budgets
Nova Pro has a reasoning mode and it is off by default. You turn it on with inferenceConfig.reasoningConfig, setting type to enabled, and you control depth with maxReasoningEffort at low, medium or high. AWS restricts this to Nova Pro and Nova Lite within Version 1. The budget is a three-level dial, not a token count, so you cannot say "spend at most 2,000 thinking tokens" the way you can on some competitors. Reasoning is visible: with ConverseStream, low and medium stream reasoning content as it is generated, while high uses different internal approaches and emits the whole reasoning trace in a final chunk - so a UI built against low will look broken at high. On Converse, reasoningConfig must go inside additionalModelRequestFields. Whether reasoning tokens bill at the output rate could not be confirmed from a primary source as of 2026-07-25; the AWS Price List carries no separate reasoning-token SKU for Nova Pro.
ToolsFunction calling and server tools
Tool calling uses toolConfig, which follows the standard Bedrock ToolConfiguration schema - so a tool definition written for Claude on Bedrock works against Nova Pro unchanged. Each toolSpec carries a name of at most 64 characters, a description and a JSON Schema inputSchema. toolChoice takes auto, any (must call something) or tool with an explicit name. Amazon documents no built-in server-side tools for Version 1; web grounding and a code interpreter appear in the Nova 2 line, not here. There is no response_format or JSON-schema output field in the Nova request schema, so structured output is done by defining a tool whose schema is your output shape and forcing it with toolChoice. Two failure modes recur. First, silently dropped parameters when topK or reasoningConfig are placed outside additionalModelRequestFields on a Converse call. Second, client timeouts: AWS allows 60 minutes per inference call while the default boto3 read timeout is 60 seconds, so long tool loops die client-side unless you raise read_timeout.
CostPrice, caching, batching, what drives the bill
Prices below are read from the AWS Price List for AmazonBedrock, us-east-1, publication version 20260723233555, which is the machine-readable source behind the Amazon Bedrock pricing page. AWS quotes per 1K tokens; the per-million column is arithmetic. What actually drives a Nova Pro bill is caching and service tier, not the headline rate. Cache writes are free and cache reads run at a quarter of the input price, so a stable system prompt or repeated document prefix is nearly free to re-send - subject to the 20K token cache ceiling, which is small enough that you must choose what to cache. Batch and Flex both halve input and output. Priority costs 75% more. Amazon does not publish a tokenizer, so token-per-word density cannot be compared with other vendors from a primary source.
| Tier | Input per 1K | Output per 1K | Per million |
|---|---|---|---|
| Standard on-demand | $0.0008 | $0.0032 | $0.80 in / $3.20 out |
| Cache read | $0.0002 | n/a | $0.20 in |
| Cache write | $0.0000 | n/a | free |
| Batch | $0.0004 | $0.0016 | $0.40 / $1.60 |
| Flex | $0.0004 | $0.0016 | $0.40 / $1.60 |
| Priority | $0.0014 | $0.0056 | $1.40 / $5.60 |
| Latency optimised | $0.0010 | $0.0040 | $1.00 / $4.00 |
| Provisioned throughput, no commit | $60.50 per model unit per hour | ||
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Amazon Bedrock | Yes | amazon.nova-pro-v1:0. Geo inference IDs us.amazon.nova-pro-v1:0 and eu.amazon.nova-pro-v1:0. Global cross-region not supported. Invoke and Converse APIs, streaming, batch. |
| A direct Amazon Nova API | No | The Nova user guide states you must access the models through Amazon Bedrock. |
| Amazon SageMaker | Unverified | Nova Version 1 customisation is documented on Bedrock. SageMaker AI hosting of Nova Pro v1 could not be confirmed from a primary source as of 2026-07-25. |
| Google Vertex AI | No | Amazon does not licence Nova to other clouds. |
| Microsoft Foundry / Azure | No | Same. |
| Hugging Face Inference | No | No weights, no provider listing. |
| Self-hosting | No | Weights are not released under any licence. |
StrengthsWhat it is good at
- 300K context with native document handling for PDF, DOCX, XLSX, HTML, CSV and Markdown, plus image and video input - documented on the Nova user guide specification table, not inferred.
- Cache writes cost nothing and cache reads cost a quarter of the input rate, so repeated prefixes are close to free within the 20K token cache ceiling.
- Four pricing tiers on the same model ID - Standard, Flex, Priority and latency-optimised - let you trade latency for a 2x price swing without changing code.
- Reasoning effort is a genuine three-level control on Nova Pro and Nova Lite, and the trace is streamable at low and medium effort.
- Everything AWS-shaped works out of the box: IAM, VPC endpoints, Guardrails, Knowledge Bases, Agents, provisioned throughput and fine-tuning.
LimitsWhere it falls down
- The output cap is 5K tokens, and AWS's own pages disagree with each other about it - the Nova user guide table says 10k while the model card and request schema say 5K. Do not design around long single-shot outputs.
- Nothing about the architecture is disclosed. No parameter count, no dense-versus-MoE, no tokenizer, no training details, no paper. You cannot reason about its behaviour from first principles.
- No
seed, nologprobs, non, no penalties and noresponse_format. Determinism and structured output both have to be faked through temperature floors and forced tool calls. - Bedrock-only, and not in every region in-region - Nova Pro is in-region in us-east-1, eu-west-2, ap-southeast-2, ap-southeast-3, me-central-1 and GovCloud, with everything else via geo cross-region inference.
- The 20K prompt-cache ceiling is small relative to the 300K context, so the long-document case that motivates the context window is largely uncacheable.
Against its neighboursHow it compares
The nearest comparison is Claude Sonnet 5, which sits behind the same Bedrock endpoint with the same tool schema - swapping is a one-line model ID change, which makes Nova Pro easy to try and easy to abandon. Nova Pro's argument against it is price and a 300K window; the argument for Sonnet is reasoning depth on hard multi-step work. Against Cohere's Command A, the comparison is different in kind: Command A gives you inline citations and a grounding documents array as first-class API features and non-commercial weights you can host, where Nova Pro gives you AWS-native operations and a cheaper rate. The other model to weigh is Amazon's own Nova 2.0 Pro, which is live in the AWS Price List at $1.375 per million input and $11 per million output - considerably more capable-priced, and where Amazon's extended-thinking and built-in-tool work is going.
Getting startedThe smallest call that works
import json
import boto3
client = boto3.client('bedrock-runtime', region_name='us-east-1')
response = client.invoke_model(
modelId='amazon.nova-pro-v1:0',
body=json.dumps({
'system': [{'text': 'Answer in one sentence. No preamble.'}],
'messages': [{
'role': 'user',
'content': [{'text': 'Summarise the termination clause.'}]
}],
'inferenceConfig': {
'maxTokens': 1024,
'temperature': 0.00001
}
})
)
print(json.loads(response['body'].read()))Change two things first. Raise the boto3 read_timeout to 3600 - AWS allows 60 minutes per call while the SDK default gives up after 60 seconds. Then add reasoningConfig with type set to enabled; it is off by default, and on Converse it must be nested inside additionalModelRequestFields or it is ignored.
SourcesWhere every claim above came from
- Nova Pro - Amazon Bedrock model card - launch date, 300K context, 5K max output, Oct 2024 cutoff, modalities, prompt caching limits, model IDs, service tiers, regional availability, sample code.
- What is Amazon Nova? - Amazon Nova user guide - family table, model IDs, document formats, language counts, distillation relationships, and the 10k max-output figure that conflicts with the model card.
- Complete request schema - Amazon Nova user guide - every inferenceConfig field, its range and default, reasoningConfig, toolConfig and toolChoice, and the 60-minute timeout guidance.
- Amazon Nova models - Amazon Bedrock inference parameters - Invoke versus Converse request patterns.
- What's new in Amazon Nova 2 - the Nova 2 line, extended thinking, built-in web grounding and code interpreter.
- AWS Price List, AmazonBedrock, us-east-1, version 20260723233555 - every price in the cost table, read directly from the AWS pricing feed because the HTML pricing page renders its tables client-side.
- Could not confirm: any architecture or parameter-count figure for Nova Pro; whether reasoning tokens are billed at the output rate; SageMaker availability for Nova Pro v1; the exact JSON key for a prompt cache checkpoint, which the Nova request schema page does not show. AWS's own two pages disagree on max output tokens and on the
topKdefault, and both disagreements are reported above rather than resolved.
What changedWhat changed here
- GPT‑6 Astra
OpenAI is rolling GPT-6 Astra out to a limited set of organizations now and to ChatGPT Plus, Pro, Business, and Enterprise users plus the API and AWS over the coming days, at API prices of $10 per million input and $50 per million output. If you are paying OpenAI rates today, this is the model to benchmark agent and computer-use work against.
Three kinds of claim, strongest first. Signal runs every morning.