Where Does Your Data Actually Live? Swiss-Hosted AI Agents vs. US Cloud LLMs
A technical comparison of Swiss-hosted, EU-hosted and US-hosted AI architectures: control, capability, latency and cost. No marketing.
There are four realistic architectures for a Swiss company: a frontier model API in a US region, the same API pinned to an EU region with zero-retention terms, an open-weight model hosted in a Swiss data centre, and a fully on-premise deployment. Control increases and capability decreases as you move down that list. Most Swiss SMEs land on the second or third option, and the correct choice is decided by data classification, not by preference.
Published 7 August 2026 · Last updated 7 August 2026
Legal note. Informational only, not legal advice. Have Swiss counsel review before publication.
The four architectures, compared
The four options differ on seven properties. Read the table by column to compare architectures, by row to find the property your constraint sits in. Data residency is a contractual and technical statement of where data is stored and processed, and it exists to give your company a defensible answer about which jurisdiction can compel access.
| Property | 1. Frontier API, US region | 2. Frontier API, EU region, zero retention | 3. Open-weight model, Swiss data centre | 4. Fully on-premise |
|---|---|---|---|---|
| Data location | US region of the provider's cloud; sub-processors may sit elsewhere | Named EU region for inference and storage; provider's home jurisdiction unchanged | Swiss data centre you choose; weights, prompts and logs under your contract | Your own premises; nothing crosses the network boundary |
| Who can technically access | Provider staff and its sub-processors; you rely on attestations | Same set, narrowed by contract; support and abuse-monitoring paths remain | Your engineers and the data centre's staff; no model vendor in the path | Your engineers only, under your own site controls |
| Model capability | Frontier: strongest on multi-step reasoning, tool use and long context | Frontier, usually the same models; new features can reach Europe later | Open-weight: adequate for extraction, classification, summarisation; weaker on long chains | As column 3, often below it: on-site hardware bounds model size |
| Latency from Switzerland | Transatlantic round trip added to generation time; noticeable in voice | Shorter European path; generation time still dominates | Shortest path; total depends on GPU sizing, not distance | Shortest path, with no elastic capacity at peak |
| Cost profile | Per token, no fixed floor; the bill follows usage | Per token, at or above the equivalent US list price | Reserved GPU capacity billed whether used or not, plus engineering time | Capital purchase, then power, cooling, refresh cycle, out-of-hours cover |
| Operational burden | Integration and monitoring only | Integration, plus contract and region verification at each renewal | Permanent: capacity, model updates, evaluation, on-call | All of column 3, plus physical infrastructure |
| Audit posture | Contracts, certifications and your own logs; the provider's estate is not inspectable | As column 1, with a documented region and retention position | You can show where the machine stands, who touched it, which version answered | Strongest evidence of control; weakest of resilience unless funded |
Do not read the table as a ranking: column 4 is not the best answer and column 1 is not the worst. A company with no personal data in scope buys only cost and delay by moving down it.
Start with data classification, not with the model
The architecture question cannot be answered until you know what the agent will read, so classify first. Three tiers are enough, and most teams finish the exercise in two hours per system.
- Public. Content you already publish or would: product descriptions, documentation, marketing copy, job adverts. No residency constraint applies; it is already outside your walls.
- Internal. Business content with no personal or client-identifying data: internal processes, non-sensitive code, meeting notes without names, aggregate figures. Confidentiality matters commercially; the regulatory surface is small.
- Regulated or personal. Anything identifying a person, or covered by professional secrecy or a client contract: names, contact details, AHV numbers, account and policy identifiers, health or employment detail. This tier drives the decision on its own.
The uncomfortable finding is usually that tier 3 is small. Most of what an internal assistant handles — drafting, translating, summarising a process — is tier 1 or 2, and routing all of it through the most restrictive architecture buys no protection while multiplying cost and latency. Swiss teams over-classify for a rational reason: under the revised Federal Act on Data Protection, fines of up to CHF 250,000 are directed at responsible natural persons rather than the company, according to Pestalozzi Attorneys at Law (2023). A written classification replaces that instinct with a decision someone else can check; the wider obligations are set out in our article on revFADP and EU AI Act compliance.
Architecture 1: frontier API, US region
This is the default configuration of nearly every AI product on the market, and for tier 1 and tier 2 content it is defensible rather than a compromise.
What you get
The strongest available models on the day they ship, with no infrastructure to run and no capacity to reserve. Tooling maturity is the underrated part: function calling, structured output, evaluation harnesses and observability all work here first. Cost scales with usage, which suits an unpredictable pilot.
What you give up
Direct control over processing location, and over who inside the provider can technically reach a request. Your assurance is contractual rather than physical: certifications, a processing agreement, a sub-processor list, audit reports. For tier 3 data a Swiss reviewer will ask you to evidence every link, and the work recurs at each renewal.
When it is genuinely fine
When no tier 3 data reaches the model and you can demonstrate that technically rather than by policy. Drafting, translation, code without secrets and internal documentation qualify. The demonstration is the hard part: an input filter, no access to customer tables, and logs proving what was sent. "Staff have been told not to paste client data" is not a control.
Architecture 2: frontier API, EU region with zero retention
This is where most Swiss SMEs with moderate sensitivity end up, and the option most often misunderstood in vendor conversations.
What the contractual terms actually cover
Zero data retention is a contractual term under which the provider does not persist prompts or outputs once a request has been served, and its purpose is to limit how long your content exists on the provider's side. Region pinning is a separate commitment covering where inference runs and where data rests, which narrows a different risk: processing in a jurisdiction you did not choose.
What they do not cover
They do not make the provider Swiss, and they do not mean your data never leaves Switzerland — it leaves the moment the request does; processing abroad is what the arrangement permits. Exceptions typically remain for abuse monitoring, for support access, and for security investigations. Zero retention at the endpoint also says nothing about retention in your own logging, often the larger exposure.
How to verify rather than trust
Ask for four documents and one test. The documents: the processing agreement naming the region, the sub-processor list with the notice period, the retention position per data flow including support, and the most recent audit report. The test: send a canary record through the production path and confirm from your egress logs where it went. Repeat at renewal; terms move and defaults re-enable.
Architecture 3: open-weight model in a Swiss data centre
This architecture answers the residency question completely, and charges for it in capability and operations. An open-weight model is a language model whose trained parameters are published for download, which allows an organisation to run inference on infrastructure it controls instead of sending requests to a vendor.
The current capability gap, stated honestly
The gap has narrowed and has not closed. On the work most business agents do — extraction, classification, summarisation, structured output, retrieval-grounded answers — a well-chosen open-weight model is usually adequate. On long autonomous chains of tool calls, on very long context and on ambiguous reasoning, frontier models remain ahead. Public leaderboards will not settle it: build twenty to fifty examples of your own task.
Real operational cost: GPUs, updates, evaluation
Three drivers dominate and only one is hardware. GPU capacity is reserved and billed whether or not you use it, sized for peak, doubled for redundancy. Model updates arrive every few months, each requiring a regression run against your evaluation set. Ownership is the third: a fraction of an engineer, permanently. The monthly Swiss infrastructure figure falls in the range of [RANGE TO BE CONFIRMED], and it is rarely the largest line.
When the sovereignty premium is worth it
When tier 3 data is central rather than incidental, when a professional secrecy duty or a client contract requires Swiss processing, or when the contractual route is expensive to evidence. High steady volume also favours it: reserved capacity beats per-token pricing above a predictable threshold. It is not worth it to reassure a board that has not read the classification.
Architecture 4: fully on-premise. Who is it for?
For a narrow set of organisations, and for most Swiss SMEs it is over-engineering. Running inference on your own hardware adds nothing to the jurisdictional answer a Swiss data centre already gives you, while adding a server room, power and cooling, physical security, a hardware refresh cycle and a resilience problem you now own alone.
The genuine cases are specific: air-gapped environments where no external network path is permitted, mandates that prohibit third-party hosting in explicit terms, GPU capacity already bought and under-used, or sites where the data cannot be moved for reasons unrelated to AI. If none of those describes you, column 3 gives the same control at a lower total cost.
One point surfaces late. The model you serve is bounded by hardware bought at one moment rather than by what has since been released, so price the same workload as Swiss colocation over three years, including the engineer, before committing capital.
The hybrid pattern that works in practice
The pattern that survives production routes each request by data class instead of choosing one architecture for everything. A rule set inspects the request, tier 1 and tier 2 traffic goes to the frontier model, and tier 3 traffic either stays on the Swiss-hosted model or passes through a redaction layer first.
The redaction layer does the work, so specify it precisely. Deterministic detection comes first: structured identifiers such as IBANs, AHV numbers and policy references match on format and should never be left to a model; named-entity recognition catches people, employers and addresses in free text. Each value is replaced by a token, with the mapping held in a vault on your side so the response can be re-populated. Pseudonymisation is the replacement of direct identifiers with reversible tokens, used to keep a record usable for processing while removing the fields that name a person.
State the limits: free text re-identifies people through combinations no detector catches — a role, a location and a date can be enough in a small market. Where the task is deterministic — moving a field, raising a ticket, reconciling two systems — process automation for Swiss SMEs solves it with no model in the path. Route to a model only what genuinely needs language.
What must FINMA-supervised institutions consider additionally?
Supervised institutions carry their existing governance, outsourcing and operational risk obligations into the deployment; the architecture choice suspends none of them. Adoption in the sector is already broad: around 50% of Swiss financial institutions use AI or have applications in development and a further 25% plan to within three years, 91% of those using AI use generative AI, institutions average five applications in production and nine in development, and the top identified risk is data quality, followed by data protection and explainability, according to a FINMA survey of roughly 400 institutions published in 2025. Two caveats before that reaches a board paper: it describes supervised financial institutions, not Swiss companies generally, and the risk ranking is self-identified.
That top-ranked risk is not a hosting question. If the source records are inconsistent, the agent will be confidently wrong in a Swiss data centre exactly as it would be in Virginia.
On supervisory expectations this article points rather than paraphrases: FINMA published Guidance 08/2024 on governance and risk management in connection with artificial intelligence, and institutions in scope should read it directly, with counsel. Note also that there is no horizontal Swiss AI statute to certify against: the Federal Council adopted a sector-specific approach to AI regulation on 12 February 2025 rather than a horizontal law, according to the Swiss Federal Council media release published on admin.ch in 2025.
How do you document the decision so it survives an audit?
Write a decision record before you build, not after the question is asked. A decision record is a short written document that states what was chosen, what was rejected and why, and it exists so that a reviewer two years later can reconstruct the reasoning without interviewing people who have left.
☐ 1. Data classes the agent will read, listed per source system and per field ☐ 2. The architecture chosen, named precisely: provider, model, region, data centre ☐ 3. The alternatives rejected, each with the reason in one sentence ☐ 4. Contractual position: processing agreement, retention terms, sub-processor list, date obtained ☐ 5. The redaction and routing rules applied, and what they do not cover ☐ 6. Retention period for prompts, outputs and logs on your own side ☐ 7. Who approved the decision, by name and role ☐ 8. Evidence held per request: model version, route taken, data class, redaction applied ☐ 9. Triggers for revisiting: new data class, new use case, contract change ☐ 10. A fixed review date, in the calendar of a named person
Point 8 is what auditors test: a log line naming route, model version and redaction outcome answers "where did this conversation go"; a monthly summary does not. Add it at build time.
Frequently asked questions
Is a Swiss-hosted model less capable than GPT or Claude?
Usually somewhat, and it depends on the task. For extraction, classification, summarisation, structured output and retrieval-grounded answers, a well-chosen open-weight model hosted in Switzerland is generally adequate. The gap shows on long chains of tool calls, on very long context and on ambiguous reasoning. Do not settle it with leaderboards: build twenty to fifty examples of your task and measure both.
Does zero data retention mean my data never leaves Switzerland?
No. Retention and processing location are separate commitments. Zero retention means the provider does not store prompts and outputs after serving the request; it says nothing about where processing happened, and the data leaves Switzerland the moment the request does. Exceptions for abuse monitoring and support access usually remain. If data must stay in Switzerland, the answer is Swiss hosting, not a retention clause.
What does sovereign hosting actually cost?
Ask for the drivers rather than one figure. Three dominate: reserved GPU capacity, sized for peak and billed whether you use it; model updates every few months, each needing a regression run against your evaluation set; and continuous ownership — evaluation, monitoring and on-call. For Swiss deployments the monthly infrastructure range is [RANGE TO BE CONFIRMED]; engineering time usually exceeds it.
Can we redact personal data before it reaches the model?
Yes, and it is the most useful single control in a hybrid design. Detect structured identifiers deterministically by format, catch names and addresses with named-entity recognition, replace each value with a token, and keep the mapping in a vault on your side so responses can be re-populated. But free text re-identifies people through combinations no detector catches, so redaction reduces risk without replacing classification.
How do we prove to an auditor where our data went?
With per-request logging and a decision record, neither of which can be reconstructed afterwards. Every request should leave a log line naming the route, the model version, the data class, whether redaction was applied and the retention period. The decision record states what was chosen, what was rejected and who approved it. A monthly summary does not answer a question about one conversation.
Do we need this if we are not a regulated firm?
Often not, and saying so matters. If your agent handles tier 1 and tier 2 content — drafting, translation, internal documentation — a frontier API under enterprise terms is a reasonable choice, and moving down the list buys cost and latency rather than protection. Sovereign hosting earns its premium when personal or confidential data is central, or when a contract or supervisor requires Swiss processing.
DINOLABS is a Colombian company that builds websites, process automation and AI agents for businesses in Colombia, Mexico, the United States and Switzerland.
To have your data classes mapped against these four architectures before you commit to one, Request a Confidential Assessment — NDA available.