Google iconGoogleAug 18, 2026 ~6 min source read

Box and Google Cloud add Gemini Multimodal Embeddings 2 to power visual, spatial, and structured enterprise agents

Box integrates Gemini Multimodal Embeddings 2 into its Agentic Platform to extend retrieval-and-agent workflows beyond text, preserving table structure, page layout, and visual content across mixed file types.

How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2

Share this story

Send the public story page.

Useful takeaways from this story.

Multimodal embeddings embed text, images, rendered pages, tables, and charts into a single semantic space so agents can retrieve visual and spatial elements as easily as prose.

Preserving layout and table semantics prevents loss of meaning that happens when converting structured data to plain text, enabling accurate financial and analytical queries.

Cross-format retrieval lets an agent link a PDF policy page, a spreadsheet table, and a presentation slide without manual tagging or format-specific workarounds.

# What changed

# Why it matters

# Core capabilities introduced

  • Unified multimodal vector space: text, raster images, rendered document pages, spreadsheet tables, and charts are embedded into the same semantic space for consistent retrieval.
  • Crossmodal retrieval: natural language queries can retrieve visual components (text-to-visual) and visual inputs can retrieve relevant text (visual-to-text) without manual tagging.
  • Layout-aware document embedding: pages are embedded as rendered images, keeping callouts, multi-column layouts, and spatial hierarchies intact instead of breaking content into arbitrary text blocks.
  • Heterogeneous format bridging: the system supports.docx,.xlsx,.pdf,.pptx,.png, and.csv while preserving modality-specific structure.

# How this changes agent behavior

# Three design patterns where multimodal agents add value 1) Complex financial and analytical reporting The embedding of rendered tables and charts preserves row/column relationships so agents can surface exact supporting tables or flag discrepancies between written claims and visual trends.

2) Multimodal clinical decision support

3) Cross-format operational workflows

# Practical effects for teams

  • Faster, more accurate retrieval of visual evidence and structured data.
  • Reduced need for manual tagging or format conversions when preparing corpora for agent use.

# What to watch next Adoption will depend on how teams pipeline existing Box content into layout-aware embeddings and how agents are trained to interpret those multimodal vectors for tasks like trend detection, anomaly identification, and document-level sourcing.

# Short summary Box plus Gemini Multimodal Embeddings 2 moves enterprise agents beyond text by embedding images, pages, tables, and charts into a unified vector space. That preserves spatial and structural meaning across mixed file types and enables agents to retrieve and reason over the precise visual or structured element that supports a query.

More context around this story.

Now introducing Gemini Enterprise for Legal
Google iconGoogleAug 25, 2026

Now introducing Gemini Enterprise for Legal

Few professions are as exacting as the practice of law. A team reviewing a contract or building a case works inside strictly privileged information, firm-specific playbooks, and a body of law that changes constantly. The work thrives on nuanced, professional judgment — and the systems supporting it inherit real obligat

Now introducing Gemini Enterprise for Financial Services
Google iconGoogleAug 25, 2026

Now introducing Gemini Enterprise for Financial Services

Protecting capital in today's markets requires immense speed and precision. A financial analyst preparing a deal memo works across licensed market data, internal models, and confidential client files. General-purpose AI lacks the real-time accuracy, verifiable data lineage, and strict security that financial institutio

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app