IT Brief Canada - Technology news for CIOs & IT decision-makers
Canada
Google Cloud & Box add Gemini to multimodal search

Google Cloud & Box add Gemini to multimodal search

Tue, 18th Aug 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

Google Cloud and Box have integrated Gemini Multimodal Embeddings 2 into Box's Agentic Platform, extending Box's content management tools beyond text-based retrieval.

The integration is designed to help Box analyse and retrieve information from documents that combine text, tables, charts, images and other visual elements. That includes PDFs, spreadsheets, slide decks, images and CSV records, which often contain business information in formats that text-only systems struggle to interpret accurately.

Enterprise content platforms have long relied on search and retrieval systems built mainly around text. That approach works for narrative material, but it can lose meaning when information depends on layout, structure or visual context, such as the relationship between a column heading and a figure in a table or the sequence shown in a process diagram.

Gemini Multimodal Embeddings 2 creates a shared semantic representation for different kinds of content, including text, raster images, document pages, rendered spreadsheet tables and visual charts. In practice, that means a user could search for a chart or diagram using natural language, or compare information across several file types without first converting everything into flat text.

Three use cases

The companies outlined three main areas where they expect the integration to be used. The first is financial and analytical reporting, where teams often work with documents containing tables, growth charts and footnotes that need to be read together rather than separately.

In those settings, a multimodal system can preserve a table's structure and connect written commentary to trends shown in charts. It can also help users locate the exact page, table or visual reference that supports a given metric in a larger document set.

The second area is healthcare and clinical work, where information may be spread across text records, photos, pathology imagery and risk matrices. According to the system description, indexing these materials in a shared representation could help connect visual findings with written medical information and triage frameworks.

The third use case is cross-document reconciliation. Businesses often need to compare content across minutes, spreadsheets, images, presentations and email records to verify details or identify inconsistencies. The integration is intended to help flag contradictions, such as outdated pricing in an image compared with a more recent spreadsheet, or differences between a scanned contract and a related legal review note.

Content shift

The development reflects a broader shift in enterprise software as companies try to make AI systems work with the full range of formats used in day-to-day operations. While retrieval-augmented generation has become a standard way to pull text from company repositories into AI tools, suppliers are now under pressure to show they can handle visual and structured information with similar accuracy.

That matters in sectors such as financial services, life sciences and legal operations, where key evidence often sits in diagrams, tables, scanned forms or slide presentations rather than plain-text documents. In those environments, the ability to preserve layout and connect related information across formats can affect how quickly staff can review records, investigate discrepancies or prepare reports.

Box has positioned the integration as part of a broader effort to turn enterprise content stores into a governed base for AI agents that search, analyse and act on company information. The approach could help organisations surface issues such as stale pricing, expiring contract clauses and contradictions between related files.

For Google Cloud, the partnership adds another example of how its Gemini models are being embedded into third-party enterprise software rather than used only through Google's own products. The company has been expanding its push into business AI tools that can work across multiple content types and connect to existing repositories of corporate data.

The technical feature at the centre of the integration is layout-aware document embedding. Instead of dividing files into arbitrary text segments, the system can embed rendered document pages directly, preserving hierarchies, callout boxes and other visual structures that often carry meaning. It also supports cross-modal retrieval, allowing text queries to return visual results and visual content to be linked back to text-based material.

Another part of the design is support for mixed-file environments. Businesses rarely keep workflows in a single format, and one process may involve a policy document, a spreadsheet log, a presentation deck and a scanned image. The shared representation is intended to bridge those sources without stripping out structure that may be essential to interpretation.

The integration underscores how enterprise AI vendors are trying to move beyond basic search to systems that can reason across a company's stored material while preserving more context. In content-heavy industries, much of that challenge comes down to whether software can read a document more like a human reviewer would, rather than treating every page as a sequence of extracted words.

Gemini Multimodal Embeddings 2 supports content including text, images, document pages, spreadsheet tables, charts and .docx, .xlsx, .pdf, .pptx, .png and .csv files.