The "Boston Model": Responsible Experimentation in Generative AI

By Idoia Ortiz de Artiñano

Co-founder and CEO of Gobe

Por Pamela Subizar

Communications Expert

Por

Por

Fecha de publicación
5/2/26
Compartir

The "Boston Model": Responsible Experimentation in Generative AI

In early 2023, Santiago Garcés, Chief Information Officer for the City of Boston, spearheaded one of the first guidelines for the responsible use of AI in local government. Since then, he and his team have built a distinct framework for innovating with this technology. In this article, we look behind the scenes of this model, examining two concrete case studies and key takeaways for municipal leaders and public administrations.

‍

OpenAI launched ChatGPT in November 2022. By January 2023, just two months later, it had already reached 100 million users, marking the fastest growth of any technology tool in history. It arrived with staggering capabilities, but also with risks that no one could, or still can, predict with certainty.

Where there is uncertainty, there is fear. And when people are afraid, the default reaction is self-protection. Most organizations, both public and private, chose what seemed like the prudent path: banning or heavily restricting the use of these tools. A tiny minority—and virtually none in the public sector—embraced them completely unchecked. Meanwhile, individual adoption kept climbing. People were willing to take personal risks to try out this undeniably useful technology, sending the exponential adoption curve nearly vertical.

This was the landscape in early 2023 when the City of Boston decided to forge a middle path between total prohibition and complete laissez-faire. CIO Santiago Garcés recognized that large language models (LLMs) and tools built on them were here to stay. Embracing them blindly was just as risky as ignoring them entirely. What the city needed instead was an environment grounded in trust and experimentation—testing what worked and what didn't, and above all, learning by building the internal skills and capacity needed to harness generative AI's potential while mitigating risks as they emerged.

In March 2023, Mayor Michelle Wu convened experts (including academics from MIT and Northeastern University) to unpack the phenomenon, while Garcés’s team drafted the city's first guidelines for responsible generative AI use, published that May. They were among the very first guidelines issued by a major U.S. city. Since then, Garcés’s 150-person team—responsible for everything from IT infrastructure maintenance to service design and data analytics—has forged its own approach to AI-driven innovation: what we might call the "Boston Model."

Recently, Santi Garcés visited Madrid to participate in Demo Day, an event co-organized by Gobe and IE University’s Center for the Governance of Change, with support from AWS. Santi also stopped by Gobe’s new offices to share their work in greater detail and offer practical lessons and insights for Spanish public sector leaders.

Here is a summary of those conversations.

1. The Three Pillars of Generative AI in Boston

When generative AI burst onto the scene, it was difficult to tell where it provided real value and where it fell short—or to determine which dimensions of public sector work stood to benefit and where the greatest risks lay.

To avoid getting lost in the hype, Boston structured its generative AI strategy around three distinct impact areas:

  • General Productivity: Deploying AI as a daily assistant for city staff (summarizing documents, drafting emails, synthesizing complex reports, coding, or analyzing data). The goal is to leverage off-the-shelf generative AI applications to streamline everyday municipal workflows.
  • Digital Transformation of Public Services: Utilizing the underlying capabilities of LLMs to directly redesign and improve public services across the city.
  • Equity: Ensuring that the integration of AI across society does not deepen existing social or economic divides, working actively to distribute the benefits of this technology as equitably as possible.

2. Case Studies: From Theory to Practice

What makes Boston’s approach compelling is not just its strategy, but the concrete projects where this vision has materialized. Of the many examples Santiago shared with us, two stand out: one focused on transforming a core citizen service (obtaining municipal permits) and another designed to support civil servants in a complex internal workflow (drafting public procurement documentation).

Redesigning the City's Permitting System

As Garcés outlined in a recent LinkedIn post, in Boston—as in many cities—there is no single permit for "installing solar panels" or "opening a restaurant." Instead, residents face a maze of hundreds of individual building, electrical, plumbing, and safety permits. Applicants were left to navigate a confusing bureaucratic labyrinth with very little guidance. In fact, city surveys revealed that 50% of Boston residents felt obtaining permits and licenses presented a major obstacle.

To solve this, the city set out to truly understand the resident experience. Permit applications included a comment field where applicants described what they were trying to accomplish while seeking a license. Prior to LLMs, analyzing those thousands of unstructured text fragments was practically impossible. Using generative AI, however, the city's data team processed 25 years of permit feedback to uncover residents' true needs. They grouped the data into roughly 200 user-experience clusters—such as "replacing my boiler" or "building a deck"—and validated these categories alongside the Inspectional Services Department. This yielded a much clearer taxonomy of what residents actually need when they first approach local government.

With a refined diagnosis in hand, the service design team partnered with Northeastern University to redesign user forms for maximum clarity, again leveraging AI tools. As a result, the City of Boston can now offer a vastly more intuitive, resident-friendly permitting experience.

BitBot: Accelerating Public Procurement

Public servants spend a significant portion of their time drafting procurement documents, but the governing regulations are notoriously complex. In Boston alone, the manual summarizing procurement rules and practices runs about 200 pages. To assist staff with this task, the Department of Innovation and Technology built BitBot: a generative AI assistant specifically tailored for city procurement workflows. Garcés's team trained BitBot on historical procurement files, state laws, local ordinances, and municipal best-practice guides. With BitBot, staff can draft requests for proposals (RFPs) and tender documents much faster while instantly resolving complex regulatory questions, slashing administrative burden while elevating document quality.

The city is currently collaborating with Harvard University to rigorously evaluate BitBot's effectiveness. Using a randomized controlled trial (RCT) design, researchers evaluated performance metrics—speed and drafting quality—between city employees using BitBot and a control group working without it. While the full academic paper is forthcoming, Santi Garcés shared at Gobe that BitBot achieved substantial gains in both drafting speed and document quality. The results were so compelling that city leadership has already decided to expand BitBot to additional municipal departments.

3. The Secret Behind the Success

While Boston's strategy and case studies are inspiring, we wanted to dig deeper. Many public administrations want to be more than just passive observers of emerging technology; they want to stay ahead of the curve, experimenting and innovating responsibly.

Therefore, we asked the Boston CIO to explain to us what capacities and ways of working with suppliers they have developed in the city. We wanted to understand What is behind the Boston model and what concrete lessons can other municipalities or administrations draw from that experience. Santi Garcés highlighted two key elements to understand their way of working and what they have achieved so far.

On the one hand, in terms of internal organization and capacity, Boston made a deliberate decision not to rely exclusively on external vendors. They have actively invested in hiring product managers, service designers, data engineers, and analysts. Organizationally, they centralized core capabilities—such as cybersecurity and emerging technologies—to serve all city departments, preventing each agency from having to "reinvent the wheel" on its own. They pair these centralized resources with small, agile teams to build tools, experiment with them, and measure their impact.

Regarding their strategy for navigating the vendor market in generative AI, Boston prioritized its autonomy. In a market where technology and business models evolve so rapidly, the city wants to avoid over-relying on a single vendor or model. Major AI companies like OpenAI and Anthropic are positioning themselves aggressively as key providers in the enterprise sector, either directly through their own products or as the engine powering third-party applications. According to recent analyses, large companies across finance, technology, industry, and professional services spent an average of $10 million on generative AI in 2025. Any organization, public or private, must understand what dependencies its approach creates with vendors if these technologies become indispensable. To safeguard its independence, rather than buying blanket commercial licenses from these providers (which run about $20 per user/month), Boston developed its own middleware platform, called "Launchpad," on a modest budget of just $10,000 per year. This platform allows them to switch AI models (OpenAI, Anthropic, Llama) as needed, paying strictly for actual API consumption rather than fixed licenses—optimizing costs while maintaining full freedom to choose the right model for the city's needs.

Innovate or Ignore? Everything Carries Risk

The Boston Model isn't perfect, nor is it a simple copy-and-paste template. Santi Garcés was refreshingly transparent about the challenges of scaling this framework organization-wide. He also acknowledged that Boston benefits from an exceptional local ecosystem of tech talent and top universities.

However, its core principles—investing in internal capabilities to continuously experiment and evaluate—hold true for any public administration. You don't need massive resources to establish guidelines for responsible use or to measure technology's real-world impact.

In fact, inaction may well be far riskier than experimentation. Public sector bodies that ignore AI aren't avoiding danger; they are simply exposing themselves to unmonitored "shadow AI" use by employees and an inability to properly assess vendors, creating hidden risks and dependencies that are hard to measure. In a period of rapid change, experimenting and innovating responsibly—by investing in internal capabilities—may actually be the safest strategy of all.

‍
‍

— This article summarizes key takeaways from Santiago Garcés's visit to Gobe in January. In addition, we have produced a podcast featuring his interview and a comprehensive report detailing the Boston Model for our subscribers. To request a copy, contact us at ventures@gobe.studio.

‍

Tech and data
GovTech Methodology

Get the best content on public digital transformation and govtech in Spanish.

Thank you so much for subscribing!
Something went wrong, please contact us by another means.