# Google DeepMind > Directory map of Google DeepMind artificial intelligence models, scientific research platforms, developer tools, and prompt configuration guides. This domain serves as the primary web property for Google DeepMind, balancing advanced scientific research publications with public and developer-facing artificial intelligence deployment models. ## Models Central directory and technical portfolio of Google DeepMind's artificial intelligence models, spanning generative media, multimodal foundation models, and scientific research systems. - [Gemini foundation model series](https://deepmind.google/models/gemini): Core foundation model series engineered with native cross-modality processing for text, images, video, audio, and code workflows. Structured into optimized execution variants including the Flash model for low-latency agentic orchestration, the Pro tier for complex logic and strategic reasoning, Deep Think for research and computational engineering, and Flash-Lite for high-throughput operational tasks. - [Gemini Pro reasoning model](https://deepmind.google/models/gemini/pro): Engineered for complex multi-step reasoning and problem-solving. Specialized for agentic performance, advanced coding pipelines, long-context processing, multimodal data ingestion, and algorithmic development. - [Gemini Flash low-latency model](https://deepmind.google/models/gemini/flash): High-throughput model configuration optimized for automated agentic workflows, programmatic coding execution, and long-horizon enterprise processes. - [Gemini Flash-Lite high-throughput model](https://deepmind.google/models/gemini/flash-lite): Scalable thinking model optimized for high-volume, cost-efficient, and low-latency tasks. Features a 1M input token context window supporting multimodal inputs (including text, image, video, audio, and PDFs) and delivers flexible reasoning levels alongside structured outputs for high-throughput orchestration pipelines. - [Gemini Deep Think reasoning mode](https://deepmind.google/models/gemini/deep-think): Specialized reasoning mode built on top of the Gemini Pro architecture, optimized for highly complex multi-step reasoning, mathematical proofing, algorithmic development, and advanced scientific problem-solving across physics, chemistry, and engineering workflows. - [Gemini multimodal embedding model](https://deepmind.google/models/gemini/embedding): Multimodal embedding model that maps text, images, video, audio, and documents into a single, unified semantic space. Features an 8,192 input token context, multilingual support for over 100 languages, and Matryoshka Representation Learning for scalable output dimensions (ranging from 128 to 3,072) to optimize downstream RAG, search, classification, and clustering workflows. - [Gemini Omni video generation amd editing model](https://deepmind.google/models/gemini-omni): Omnimodal model engineered for high-resolution video generation from multimodal inputs across diverse visual styles. Supports real-world physics simulation, conversational video editing, and instruction execution across variable prompt complexities. - [Gemini Image aka Nano Banana generative visual model](https://deepmind.google/models/gemini-image): Generative visual model optimized for high-precision image creation, image modification, and rapid rendering iterations. Specialized for embedding legible text within diagrams or posters, executing multi-language localized text rendering, and applying long-context factual knowledge to generate complex infographics and accurate historical scenes. - [Gemini Audio processing model](https://deepmind.google/models/gemini-audio): Audio processing model family designed for streaming multi-turn audio, video, and text inputs. Optimized for synchronous voice and video interactions, continuous audio stream translation, and programmatic text-to-speech (TTS) synthesis. - [Lyria AI music generation model](https://deepmind.google/models/lyria): Family of AI music generation models, including Lyria 3 and Lyria RealTime, designed for composing high-fidelity music tracks up to three minutes long. Supports multimodal image-to-audio prompting, multilingual vocal generation, detailed acoustic control, and integrated SynthID watermarking for secure content authentication. - [Veo generative video model](https://deepmind.google/models/veo): Generative video model series, featuring Veo 3.1, engineered to produce high-definition cinematic video with native, synchronized audio generation (including dialogue, ambient sound, and sound effects). Built with advanced physical world modeling and precise instruction adherence to ensure temporal consistency, realism, and creative control. - [Imagen text-to-image model](https://deepmind.google/models/imagen): Text-to-image generative model series, including Imagen 4, designed to produce high-fidelity visuals across diverse artistic styles ranging from photo-realism to abstract illustration. Supports high-resolution generation up to 2k with precise spelling and typography rendering, and features a high-efficiency mode optimized for low-latency generation workflows. - [Gemma open foundation model](https://deepmind.google/models/gemma): Open foundation model series supporting multimodal processing across language, vision, and acoustic data. Configured for programmatic text generation, document summarization, interactive conversational agents, visual data extraction, speech transcription, and exploratory research in natural language processing (NLP) and vision-language frameworks. - [Genie interactive world model](https://deepmind.google/models/genie): Interactive world model series, featuring Genie 3, designed to generate and explore controllable, photorealistic 3D environments from simple text prompts. Supports real-time rendering at 720p (operating at 20–24 fps) with stable spatial consistency for sustained interaction, providing simulated environments to train and evaluate autonomous agents. - [Gemini Robotics embodied AI suite](https://deepmind.google/models/gemini-robotics): Embodied AI model suite utilizing a dual-architecture system that pairs Vision-Language-Action (VLA) and Embodied Reasoning (ER) technologies. Includes Gemini Robotics 1.5 (VLA) for mapping visual and textual prompts directly to motor commands across diverse hardware platforms, and Gemini Robotics-ER 1.6 for physical-world planning, autonomous decision-making, and natural language coordination. - [SynthID digital watermarking technology](https://deepmind.google/models/synthid): Watermarking technology designed to embed imperceptible, tamper-resistant digital identifiers directly into AI-generated text, images, video, and audio. It adjusts token probability scores for text generation and introduces robust, invisible markers to media files, preserving content quality while enabling verification via the SynthID Detector portal or direct integration with Gemini. ## Research Central directory of Google DeepMind's scientific publications, safety evaluations, and technical milestones. Details ongoing progress in fundamental AI domains including general-purpose world modeling (Genie 3), reinforcement learning systems (AlphaGo/AlphaZero), virtual agents (SIMA 2), embodied robotics, and alignment research. - [DiLoCo distributed training architecture](https://deepmind.google/blog/decoupled-diloco): Distributed, low-communication AI training architecture designed to train large language models across geographically distant data centers. By structuring training runs across decoupled "islands" of compute with asynchronous data flow, it isolates local hardware failures to enable continuous, self-healing training, operates efficiently over standard wide-area network bandwidth (2–5 Gbps), and supports mixing heterogeneous hardware generations in a single run without sacrificing performance. - [AI Pointer context-aware interface](https://deepmind.google/blog/ai-pointer): Experimental interface framework powered by Gemini that reimagines the traditional mouse cursor as a context-aware interaction tool. By capturing the visual and semantic context surrounding the pointer, it enables users to execute shorthand, cross-application commands (such as voice/text prompts like "Fix this" or "Summarize that") and dynamically translates static screen pixels into interactive, structured digital entities. - [CodeMender automated software security agent](https://deepmind.google/blog/introducing-codemender-an-ai-agent-for-code-security/): AI-powered agent for automated software security that identifies, patches, and proactively refactors vulnerable source code. Powered by Gemini Deep Think models, it utilizes advanced program analysis (including static, dynamic, differential testing, and SMT solvers) alongside a multi-agent critique loop to generate and validate regression-free security patches and secure API rewrites for production codebases. - [SIMA 2 virtual environment agent](https://deepmind.google/blog/sima-2-an-agent-that-plays-reasons-and-learns-with-you-in-virtual-3d-worlds/): Gemini-powered AI agent designed for interactive reasoning and goal execution across diverse, virtual 3D environments. Operating via standard visual inputs and keyboard/mouse controls without direct game-engine access, it generalizes complex instructions across unseen games and newly generated world simulations (Genie 3) while demonstrating closed-loop, self-improvement capabilities via autonomous play. - [AlphaGo reinforcement learning system](https://deepmind.google/research/alphago): Historical reinforcement learning system and the first computer program to defeat a professional human Go player and a world champion. It combined deep neural networks—utilizing policy networks for move selection and value networks for position evaluation—with Monte Carlo Tree Search (MCTS) to navigate the game's vast search space. - [AlphaZero self-play reinforcement algorithm](https://deepmind.google/blog/alphazero-shedding-new-light-on-chess-shogi-and-go/): Self-taught reinforcement learning algorithm that mastered chess, shogi, and Go starting from scratch through self-play. It combines deep neural networks with Monte Carlo Tree Search (MCTS) to evaluate positions, replacing human-engineered heuristics and brute-force search databases with learned strategic intuition. - [MuZero model-based planning algorithm](https://deepmind.google/blog/muzero-mastering-go-chess-shogi-and-atari-without-rules/): Model-based reinforcement learning algorithm that learns to master complex environments—including chess, shogi, Go, and Atari—without being programmed with the underlying rules. It achieves this by modeling only the variables critical to planning and decision-making: the value of the current state, the optimal policy to take, and the immediate reward. ## Science Machine learning architectures and computational systems applied to disciplines across the natural, empirical, and mathematical sciences, including structural biology, predictive meteorology, and algorithmic discovery. - [AlphaFold protein structure prediction model](https://deepmind.google/science/alphafold): Deep learning architecture engineered to predict three-dimensional protein structures and biomolecular complexes from primary amino acid sequences. Connects to the AlphaFold Protein Structure Database and AlphaFold Server to provide structural data maps and interaction modeling tools for non-commercial scientific research. - [WeatherNext atmospheric forecasting model](https://deepmind.google/science/weathernext): Global medium-range atmospheric forecasting model family built on a Functional Generative Network (FGN) architecture. Computes 15-day global trajectory ensembles on a 0.25° spatial grid with up to 1-hour temporal resolution, predicting core atmospheric and surface variables to track low-probability extreme weather events and rapid intensification. - [Gemini for Science](https://ai.google/gemini-for-science): Ecosystem of experimental analytical tools and multi-agent frameworks optimized to automate core stages of the scientific method. - [AlphaGenome DNA sequence model](https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome): Unifying DNA sequence model that predicts how genetic variants and mutations affect regulatory activity across coding and non-coding regions. It processes long sequence context up to 1 million base-pairs at single-base resolution, mapping molecular properties such as RNA splicing, gene boundaries, and accessibility to accelerate research in disease biology and synthetic genome design. - [AlphaMissense genetic mutation pathogenicity model](https://deepmind.google/blog/a-catalogue-of-genetic-mutations-to-help-pinpoint-the-cause-of-diseases): AI model built on the AlphaFold architecture to predict the pathogenicity of missense mutations within human proteins. By analyzing protein sequence data and structural context, it classifies 89% of all 71 million possible missense variants as either likely benign or likely pathogenic to accelerate clinical genetics and rare disease research. - [AlphaProteo protein binder design system](https://deepmind.google/blog/alphaproteo-generates-novel-proteins-for-biology-and-health-research/): AI system designed to generate novel, highly specific protein binders that target and attach to designated proteins. Utilizing structural data from the Protein Data Bank and AlphaFold predictions, it automates the design of custom binder proteins to accelerate drug discovery, diagnostic development, and bioscience research. - [AlphaEarth Foundations geospatial AI model](https://deepmind.google/blog/alphaearth-foundations-helps-map-our-planet-in-unprecedented-detail): Geospatial AI foundation model designed to synthesize multi-source Earth observation data (including optical, radar, 3D laser mapping, and climate datasets) into compact 64-dimensional embeddings at a 10x10 meter resolution. Integrated into Google Earth Engine as the Satellite Embedding dataset, it enables planetary-scale tracking of environmental, agricultural, and ecological shifts with reduced computational storage overhead. - [Perch bioacoustic conservation model](https://deepmind.google/blog/how-ai-is-helping-advance-the-science-of-bioacoustics-to-save-endangered-species/): Open-source bioacoustic AI model designed to help conservationists analyze environmental audio data. Trained on extensive datasets spanning birds, mammals, amphibians, and anthropogenic noise across terrestrial and marine (coral reef) ecosystems, it supports vector search and agile modeling to rapidly build custom classifiers from a single audio sample to monitor wildlife populations. - [Fusion plasma control reinforcement system](https://deepmind.google/blog/accelerating-fusion-science-through-learned-plasma-control): Deep reinforcement learning (RL) control system designed to autonomously stabilize and manipulate superheated hydrogen plasma within a nuclear fusion tokamak. Using a single neural network to dynamically coordinate 19 magnetic coils simultaneously, it maintains plasma containment and sculpts complex experimental configurations (such as "snowflake" and double-droplet shapes) to assist in reactor design, supported by the JAX-based TORAX open-source plasma simulator. - [GNoME crystal structure prediction framework](https://deepmind.google/blog/millions-of-new-materials-discovered-with-deep-learning): Deep learning framework based on Graph Neural Networks (GNNs) designed to predict the stability and structure of inorganic crystalline materials. Powered by an active learning pipeline validated with Density Functional Theory (DFT) into its training dataset, it discovered 2.2 million new crystal structures, contributing 380,000 highly stable candidate materials to the Materials Project to assist in the experimental synthesis of superconductors, solid-state batteries, and electronics. - [Deep Loop Shaping feedback control method](https://deepmind.google/blog/using-ai-to-perceive-the-universe-in-greater-depth): Reinforcement learning control method designed to optimize high-sensitivity feedback systems in gravitational-wave observatories like LIGO. Trained using frequency-domain rewards, it suppresses active vibrations and control noise in unstable feedback loops by 30 to 100 times without amplifying adjacent frequency bands, with broader applications across robotics, aerospace, and precision mechanical engineering. - [AlphaQubit quantum error correction decoder](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/alphaqubit-quantum-error-correction/): AI-based quantum decoder designed to identify and correct physical errors in superconducting quantum processors. Utilizing deep learning architectures trained on data from Sycamore processors, it analyzes experimental measurement histories to classify error events (bit-flips and phase-flips) with high accuracy, outperforming traditional algorithmic decoders to assist in scaling stable logical qubits. - [Quantum Matter density functional simulation](https://deepmind.google/blog/simulating-matter-on-the-quantum-scale-with-ai): Neural network density functional for Density Functional Theory (DFT) designed to simulate the quantum mechanical behavior of electrons in molecules. By incorporating fractional electron behaviors into its training dataset, it corrects long-standing systematic errors in computational chemistry—specifically delocalization error and artificial spin-symmetry breaking—to improve the modeling of chemical reactions, charge transfers, and molecular bond dynamics. - [Fluid Dynamics PINN discovery framework](https://deepmind.google/blog/discovering-new-solutions-to-century-old-problems-in-fluid-dynamics): Mathematical discovery framework based on Physics-Informed Neural Networks (PINNs) designed to systematically identify unstable singularities ("blow ups") in complex fluid equations. By training neural networks directly on physical laws using second-order optimizers to achieve near-machine precision, it maps mathematical patterns in singularity speed and instability to assist in computer-assisted proofs and fluid mechanics research. - [AlphaEvolve algorithm discovery agent](https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms): Evolutionary coding agent powered by Gemini models (utilizing Gemini Flash for exploration breadth and Gemini Pro for analytical depth) designed for general-purpose algorithm discovery and optimization. By combining LLM-driven codebase generation with automated execution and verification loops, it iteratively refactors complex algorithms to optimize digital infrastructure (such as GPU kernels, TPU chip architectures, and data center scheduling) and solve open mathematical problems. - [AlphaProof mathematical reasoning model](https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad): Advanced reasoning model configuration designed for complex, multi-step problem solving, demonstrated by achieving a certified gold-medal standard score at the 2025 International Mathematical Olympiad (IMO). Operating end-to-end in natural language within a standard 4.5-hour time limit, it utilizes parallel thinking architectures to simultaneously explore and synthesize multiple solution paths, combined with novel reinforcement learning algorithms optimized for theorem-proving data. - [AlphaGeometry neuro-symbolic geometry system](https://deepmind.google/blog/alphageometry-an-olympiad-level-ai-system-for-geometry): Neuro-symbolic AI system designed to solve complex geometry problems at an Olympiad-level. By pairing a neural language model for intuitive planning (predicting auxiliary points and lines) with a symbolic deduction engine for logical proof construction, it executes rigorous geometric reasoning using a training set of 100 million synthetic geometry proofs generated without human demonstrations. - [AlphaChip reinforcement learning floorplanning system](https://deepmind.google/blog/how-alphachip-transformed-computer-chip-design): Reinforcement learning system designed to optimize computer chip floorplanning by automating the physical placement of blocks (such as memory controllers and processing cores). Treating the layout process as a sequential game to optimize for power, performance, and wire length, it generalizes across architectures to generate optimal physical layouts in hours that traditionally required weeks of human effort, and has been used to design multiple generations of Google TPUs. - [FunSearch mathematical algorithm discovery tool](https://deepmind.google/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-models): Search-based algorithm discovery framework that pairs an LLM with an automated evaluator to generate verified computer programs for complex mathematical and computer science problems. By using an evolutionary feedback loop that progressively refines a pool of functional code, it has discovered new solutions in extremal combinatorics (the cap set problem) and generated optimized heuristics for NP-hard resource allocation (the bin packing problem). - [AlphaDev assembly algorithm discovery system](https://deepmind.google/blog/alphadev-discovers-faster-sorting-algorithms): Reinforcement learning-based system designed to discover optimized computer science algorithms by operating directly at the assembly instruction level. By modeling algorithm construction as a sequential game and evaluating generated assembly instructions for correctness and execution latency, it discovered faster sorting and hashing algorithms that have been integrated into the LLVM C++ standard library. ## About - [About](https://deepmind.google/about): Overview of Google DeepMind's mission, history, and scientific philosophy. Outlines the unified research organization's structure—following the integration of Google Brain and DeepMind—and its dual focus on pioneering foundational artificial intelligence systems and responsibly applying AI to solve major scientific and societal challenges. - [Responsibility & Safety](https://deepmind.google/responsibility-and-safety): Strategic hub outlining Google DeepMind's governance frameworks, risk mitigation protocols, and ethical guardrails. Details the operations of the Responsibility and Safety Council (RSC) and the AGI Safety Council, the execution of the Frontier Safety Framework to manage potential systemic risks from powerful models, and collaborative partnerships (such as the Frontier Model Forum) aimed at advancing secure, privacy-preserving, and globally accessible AI technologies. - [National Partnerships for AI](https://deepmind.google/national-partnerships-for-ai): Global collaborative framework through which Google DeepMind partners with governments (including the US, UK, India, South Korea, and Singapore) to apply frontier AI models to public interests. It focuses on accelerating scientific research, modernizing public services, expanding localized educational resources, and conducting joint safety and resilience evaluations with national AI security institutes. - [The Podcast](https://deepmind.google/the-podcast): Audio series hosted by mathematician Hannah Fry, featuring in-depth interviews with leading researchers, engineers, and scientists. It provides behind-the-scenes insights into the development, technical challenges, and societal implications of key AI milestones, including AlphaFold, reinforcement learning systems, model alignment, and the future of artificial general intelligence.