Skip to content

Google DeepMind Gemma 4 Open-Weight Models Now Available on Amazon Bedrock

AWS added Google DeepMind's Gemma 4 31B, 26B MoE, and E2B variants to Amazon Bedrock, giving teams managed access to the Apache 2.0 licensed open-weight models with native multimodal support.

Amazon Bedrock logo icon in purple gradient on light background
Credit: AWS

Google DeepMind's Gemma 4 open-weight model family is now available on Amazon Bedrock, giving AWS customers managed access to three instruction-tuned variants without provisioning their own GPU infrastructure.

AWS announced the availability of Gemma 4 31B (dense), Gemma 4 26B-A4B (mixture-of-experts), and Gemma 4 E2B (edge-optimized) on the Bedrock platform. All three support text and image input, native function calling, and structured JSON output. The 31B and 26B models offer up to a 256K context window; the E2B is rated to 128K. According to the Artificial Analysis Intelligence Index cited in the announcement, Gemma 4 31B scores 39 against a median of 15 in the 4B-to-40B open-weights class, which AWS frames as competitive performance-per-parameter at this price point.

Google DeepMind's open-weight license and AWS infrastructure

Google DeepMind released all Gemma 4 weights under the Apache 2.0 license, so organizations can fine-tune on their own data and deploy the results across cloud or on-premises environments. On Bedrock, inference runs entirely on AWS infrastructure, which matters for customers with data-residency or regulatory constraints that prevent using a third-party model API. The Bedrock service layer adds the security and access controls AWS already provides for its other managed foundation models.

Google DeepMind introduced Gemma 4 in April across four sizes, including E2B and E4B edge variants optimized for on-device inference on Android and IoT hardware. The Bedrock landing extends the reach to teams whose primary infrastructure is AWS and who want to avoid managing model serving themselves. The 26B mixture-of-experts variant activates only 3.8 billion parameters per inference request, which AWS notes as a latency advantage for high-throughput workloads.

Teams evaluating open-weight models on Bedrock can now compare Google DeepMind's Gemma 4 directly against Meta's Llama series and other open models available in the service, without switching platforms or managing a separate serving stack.

Share this story

Julian Beaumont

Julian Beaumont covers artificial intelligence and large language models for techshooked, following the path from research paper to deployed feature. His standard is anti-hype: ask what a model actually does, what trained it, how it fails, and whether a benchmark measures what the announcement claims.