<?xml version="1.0" encoding="utf-8"?><rss version="2.0">
  <channel>
    <title>Liip Blog</title>
    <link>https://www.liip.ch/de/blog</link>
    <lastBuildDate></lastBuildDate>
            <item>
      <title>ChatGPT now speaks WebMCP, LiipGPT is ready</title>
      <link>https://www.liip.ch/de/blog/chatgpt-now-speaks-webmcp-liipgpt-is-ready</link>
      <guid>https://www.liip.ch/de/blog/chatgpt-now-speaks-webmcp-liipgpt-is-ready</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 +0200</pubDate>
      <description><![CDATA[<p>Since the end of August, ChatGPT (Codex or Work) has supported WebMCP by default. Finally, the work we did six months ago <a href="https://www.liip.ch/en/blog/webmcp-making-liipgpt-tools-discoverable-by-browser-ai-agents">integrating WebMCP into LiipGPT</a> is paying off. And it's no longer just an obscure feature flag in Chrome that enables people to use it.</p>
<p>No more adding MCP servers manually to ChatGPT, just give it a hint where to look (I assume with a more obvious or well-known domain, it would figure that out by itself). Maybe not "no more," but hopefully less setup and more discovery. Authentication is also usually solved with WebMCP in the way users are used to: it runs in a browser session, where you might need to log in to a site first. No more OAuth workflows that are obscure to the casual user.</p>
<p>Example Question:</p>
<blockquote>
<p>When is the next cardboard collection at Quellenstrasse? Check <a href="https://zuericitygpt.ch">zuericitygpt.ch</a></p>
</blockquote>
<p>Answer in ChatGPT:</p>
<blockquote>
<p>The next cardboard collection on Quellenstrasse (8005) is Wednesday, 9 September 2026. (ZüriCityGPT)</p>
</blockquote>
<p>And it didn't just hallucinate it, or somehow find it otherwise, it actually called the MCP tool <code>get_waste_collection</code>.</p>
<p>It also works for follow up questions for a different place or even to check out a City Council resolution (Stadtratsbeschluss), because ZüriCityGPT has maybe the best approach to search for them (and everything else on the City of Zurich website, but that a normal WebSearch usually also figures out).</p>
<figure><a href="https://www.liip.ch/&lt;https://liip.rokka.io/dynamic/o-af-1/84cc51/chatgpt-webmcp.png&gt;" rel="noreferrer" target="_blank"><img alt="" src="https://liip.rokka.io/www_inarticle_5/84cc51/cleanshot-2026-09-03-at-22-21-34.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/84cc51/cleanshot-2026-09-03-at-22-21-34.jpg 2x"></a></figure>
<p>And because I like the approach a lot, I also added WebMCP to our instance of the <a href="https://www.liip.ch/en/blog/city-of-zurich-s-900-open-data-sets-now-have-an-mcp-server">City of Zurich Open Data MCP Server</a> at <a href="https://zurich-opendata-mcp.liipgpt.ch/">https://zurich-opendata-mcp.liipgpt.ch/</a></p>
<p>From the example question in <a href="https://github.com/malkreide/zurich-opendata-mcp">their GitHub repo</a>: </p>
<blockquote>
<p>How many free parking spots are in Zurich right now, and what's the air quality like? Check on <a href="https://zurich-opendata-mcp.liipgpt.ch/">https://zurich-opendata-mcp.liipgpt.ch/</a></p>
</blockquote>
<p>Answer in ChatGPT:</p>
<blockquote>
<p>The service reports 6,862 free spaces across 36 parking facilities, but that figure is not genuinely current: its timestamp is 18 July 2026, 05:27, and several capacity figures are internally inconsistent. So there’s no trustworthy “right now” parking total from this feed.<br />
Air quality is generally good, based on provisional hourly readings from 4 September 2026, 06:00:</p>
<ul>
<li>PM2.5: 4.2–7.0 µg/m³</li>
<li>PM10: 8.1–16.6 µg/m³</li>
<li>Ozone: 7.9–73.7 µg/m³</li>
<li>NO₂: 4.9–49.4 µg/m³ — highest near Schimmelstrasse and Rosengartenstrasse</li>
</ul>
<p>In short: low particulate pollution, with moderately elevated traffic-related NO₂ at central roadside stations. Source: Zürich Open Data MCP.</p>
</blockquote>
<figure><a href="https://liip.rokka.io/dynamic/o-af-1/8c5d8e/cleanshot-2026-09-04-at-08-52-53.png" rel="noreferrer" target="_blank"><img alt="" src="https://liip.rokka.io/www_inarticle_5/8c5d8e/cleanshot-2026-09-04-at-08-52-53.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/8c5d8e/cleanshot-2026-09-04-at-08-52-53.jpg 2x"></a></figure>
<p>These examples are not groundbreaking, even if asking for "recycling collection times" is always a beloved topic on our public chatbots and Google also doesn't know them by heart. But if websites start adding more useful WebMCP endpoints to their websites, beyond just searching their content, it will enhance the experience for people using ChatGPT et al. Like adding items to a shopping cart or sending a report about a broken street sign or just help filling out a webform without actually sending it. Without them even knowing what WebMCP or MCP is or registering anything explicitly.</p>
<p>Very much looking forward to future use cases for WebMCP and how it will improve the experience with tools like ChatGPT (and others, when they implement it)</p>
<p>See also <a href="https://learn.chatgpt.com/docs/webmcp">OpenAI's documentation about WebMCP</a> for more about their implementation.</p>]]></description>
    </item>
        <item>
      <title>Apertus 1.5 - 6 Ways to Try Out Switzerland&#039;s Updated AI Model</title>
      <link>https://www.liip.ch/de/blog/apertus-1-5-6-ways-to-try-out-switzerland-s-updated-ai-model</link>
      <guid>https://www.liip.ch/de/blog/apertus-1-5-6-ways-to-try-out-switzerland-s-updated-ai-model</guid>
      <pubDate>Thu, 03 Sep 2026 00:00:00 +0200</pubDate>
      <description><![CDATA[<p>On July 24, 2026, the <a href="https://ai.ethz.ch/news-and-events/ai-center-news/2026/07/apertus-15-building-the-next-generation-of-open-ai-infrastructure.html">Swiss AI Initiative</a> — EPFL, ETH Zurich, and CSCS — released <a href="https://publicai.co/stories/apertus-1-5">Apertus 1.5</a>, another step forward for open, transparent, and sovereign AI. Building on Apertus 1.0, this update brings a massively expanded context window (up 4x, to 262,144 tokens), an optional "Thinking Mode" for reasoning, native image understanding (with experimental audio input), and meaningfully better instruction-following and tool use.</p>
<p>At Liip, we're glad to see this foundation model keep improving — it's what powers sovereign, ethical AI use cases for organizations in Switzerland and beyond. We <a href="https://www.liip.ch/en/blog/apertus-after-8-months-what-we-learned-and-look-forward-using-switzerlands-ai-model">wrote up our own hands-on comparison</a> after running Apertus in production on ZüriCityGPT, so if you want the deep-dive data, that's the place to look. Here, in the spirit of our <a href="https://www.liip.ch/en/blog/apertus-4-ways-to-try-out-switzerland-s-new-ai-model">original &quot;4 ways&quot; post</a>, we're keeping it practical: 6 ways to try Apertus 1.5 yourself, today.</p>
<h2>1. Try Apertus 1.5 on Public AI</h2>
<figure><a href="https://liip.rokka.io/www_inarticle_5/efe186/01-public-ai-thinking-mode.jpg"><img alt="" src="https://liip.rokka.io/www_inarticle_5/efe186/01-public-ai-thinking-mode.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/efe186/01-public-ai-thinking-mode.jpg 2x"></a><figcaption>Apertus 1.5's Thinking Mode running on Public AI</figcaption></figure>
<p>Use this when you just want to try the model in a chat window, right now, without installing anything or writing a line of code. <a href="https://publicai.co/">Public AI</a>'s mission is to make open AI accessible to everyone, and it remains one of the easiest ways to access Switzerland's sovereign models. It runs on Public AI's own inference utility, hosted within Swiss jurisdiction, and with the 1.5 release you can put the improved multilingual handling — including the well-liked Schwizerdütsch toggle — to the test yourself.</p>
<p>Create a free account at <a href="https://chat.publicai.co/">chat.publicai.co</a> to try the full 70B reasoning model, and make sure to switch on the new <strong>Thinking Mode</strong>. Prefer not to sign in? <a href="https://publicai.co/chat">publicai.co/chat</a> gives you the 8B version with no login required.</p>
<h2>2. Try Apertus 1.5 with RAG using LiipGPT</h2>
<figure><a href="https://liip.rokka.io/www_inarticle_5/4809b7/02-liipgpt-zuericitygpt-oss.jpg"><img alt="" src="https://liip.rokka.io/www_inarticle_5/4809b7/02-liipgpt-zuericitygpt-oss.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/4809b7/02-liipgpt-zuericitygpt-oss.jpg 2x"></a><figcaption>Apertus 1.5 answering a question via LiipGPT's open-source ZüriCityGPT integration</figcaption></figure>
<p>Use this when you want to see Apertus power a real Retrieval-Augmented Generation (RAG) chatbot that answers questions grounded in your own documents, rather than a bare chat window. <a href="https://liipgpt.ch">LiipGPT</a>, our generative AI platform, also runs on open-source models, and we're updating our open-source implementations to use Apertus 1.5. The new 262,144-token context window — 4x the original — is especially useful here: it lets our RAG setups pull in much larger documents, reports, and datasets at once without losing detail.</p>
<p>You can see this in action in our open-source integration: <a href="https://oss.zuericitygpt.ch/">oss.zuericitygpt.ch</a>.</p>
<h2>3. Try Apertus 1.5 in TextMate</h2>
<figure><a href="https://liip.rokka.io/www_inarticle_5/2422da/03-textmate-apertus-revision.jpg"><img alt="" src="https://liip.rokka.io/www_inarticle_5/2422da/03-textmate-apertus-revision.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/2422da/03-textmate-apertus-revision.jpg 2x"></a><figcaption>TextMate revising a text using Apertus 1.5</figcaption></figure>
<p>Use this when you want to feel the model's improvements on your own everyday writing — a report to tidy up, a text to translate — without any setup at all. <a href="https://liip-textmate.liipgpt.ch/">TextMate</a>, our LiipGPT module for text correction, plain-language rewriting, and translation across 12+ languages, now offers Apertus 1.5 as a model option too. It's a good way to get a feel for the model's improved instruction-following on everyday editorial work — correcting, rewriting, or translating your own text.</p>
<p>Try it directly at <a href="https://liip-textmate.liipgpt.ch/">liip-textmate.liipgpt.ch</a>.</p>
<h2>4. Download and test Apertus 1.5 locally</h2>
<figure><a href="https://liip.rokka.io/www_inarticle_5/df7e79/04-huggingface-apertus-v1-5-8b.jpg"><img alt="" src="https://liip.rokka.io/www_inarticle_5/df7e79/04-huggingface-apertus-v1-5-8b.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/df7e79/04-huggingface-apertus-v1-5-8b.jpg 2x"></a><figcaption>The Apertus 1.5 8B model page on Hugging Face</figcaption></figure>
<p>Use this when you need the model running fully under your own control — offline, air-gapped, or wired directly into your own application, with nothing sent to a third party. As with every Apertus release, the 1.5 model family — both the 8B and 70B versions — remains fully open: open weights, open data, open training recipes. You can download the models directly from Hugging Face (<a href="https://huggingface.co/swiss-ai/Apertus-v1.5-8B">swiss-ai/Apertus-v1.5-8B</a> and <a href="https://huggingface.co/swiss-ai/Apertus-v1.5-70B">swiss-ai/Apertus-v1.5-70B</a>). If you'd rather not manage inference yourself, Infomaniak also offers Apertus 1.5 70B through their API.</p>
<p>The 8B version is well-optimized and runs comfortably on modern hardware. On a Mac, <a href="https://github.com/ml-explore/mlx-lm">mlx-lm</a> gets you running locally via Apple Silicon's unified memory and Metal GPU acceleration; with quantization, people have gotten it working even on an <a href="https://digitalpathlines.ch/2026/08/12/shrinking-apertus-1-5-8b-for-an-8-gb-laptop-gpu/">8GB laptop GPU</a>. On other setups, <a href="https://github.com/vllm-project/vllm">vLLM</a> is a solid way to serve it and explore the native image-understanding features.</p>
<h2>5. Explore building your own chatbots with Open WebUI</h2>
<figure><a href="https://liip.rokka.io/www_inarticle_5/f659b1/05-open-webui-side-by-side.jpg"><img alt="" src="https://liip.rokka.io/www_inarticle_5/f659b1/05-open-webui-side-by-side.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/f659b1/05-open-webui-side-by-side.jpg 2x"></a><figcaption>Liip's Open WebUI instance answering the same prompt side by side on aisingapore/Gemma-SEA-LION-v4-27B-IT and swiss-ai/apertus-v1.5-70b-thinking</figcaption></figure>
<p>Use this when you want to build a custom, self-hosted chatbot for your team or organization — and compare Apertus 1.5 against other models before committing to one. <a href="https://www.liip.ch/en/open-webui">Open WebUI</a> remains a great self-hosted, open-source platform for building ChatGPT-like custom chatbot experiences. With Apertus 1.5's improved tool use and instruction-following, it's easier than ever to build bespoke assistants tailored to your organization's workflows.</p>
<p>Apertus 1.5 plugs into Open WebUI just like any other model — it's one of the models available on our own internal Open WebUI instance. One of the platform's best features is the ability to run the same prompt against several LLMs at once — handy if your team wants to quickly see how Apertus 1.5 stacks up against other models for your specific use case.</p>
<h2>6. Let Project Agent draft a ticket for you</h2>
<figure><a href="https://liip.rokka.io/www_inarticle_5/bd9b2a/06-project-agent-youtrack-ticket.jpg"><img alt="" src="https://liip.rokka.io/www_inarticle_5/bd9b2a/06-project-agent-youtrack-ticket.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/bd9b2a/06-project-agent-youtrack-ticket.jpg 2x"></a><figcaption>Project Agent (PAL) drafting a YouTrack ticket using Apertus 1.5's tool-calling</figcaption></figure>
<p>Use this when you want to see Apertus 1.5 not just answer questions but actually do something — turning a rough problem description into a properly researched, ready-to-file ticket. Our internal Project Agent (PAL) — a conversational agent platform Liip teams use on their own projects — now runs on Apertus 1.5 too, and the model's improved tool use makes a real difference here. Point it at a rough problem description, and it uses tool calls to dig into the project itself: reading the git history, searching the issue tracker, and pulling in the relevant wiki pages, then drafts a properly scoped improvement ticket — priority, affected component, concrete impact — ready to send straight to YouTrack.</p>
<p>It's a good showcase of what "better tool use" actually buys you in practice: not a chatbot answering questions about tools, but a model reliably calling the right ones, in the right order, to produce something you'd otherwise have written by hand.</p>
<hr />
<p>How do you want to put the new multimodal and reasoning capabilities of secure, sovereign large language models to work? Let's explore together how to build your own Open WebUI–based chatbot infrastructure, powered by Apertus 1.5.</p>
<p>Looking forward to hearing from you.</p>
<p>Next, we're looking forward to exploring Apertus 1.5's translation capabilities in more depth.</p>]]></description>
    </item>
        <item>
      <title>Apertus 1.5 - first impressions from using Switzerland&#8217;s updated AI model</title>
      <link>https://www.liip.ch/de/blog/apertus-1-5-first-impressions-from-using-switzerland-s-updated-ai-model</link>
      <guid>https://www.liip.ch/de/blog/apertus-1-5-first-impressions-from-using-switzerland-s-updated-ai-model</guid>
      <pubDate>Wed, 26 Aug 2026 00:00:00 +0200</pubDate>
      <description><![CDATA[<p>On July 24, the Swiss AI Initiative — EPFL, ETH Zurich, and CSCS — released Apertus 1.5. By July 25, we had already deployed it to <a href="https://oss.zuericitygpt.ch">oss.zuericitygpt.ch</a>. In this blogpost we will take a look at our first impressions testing Apertus 1.5 on the OSS version of ZueriCityGPT.</p>
<p>It was not just the open-source side that upgraded. The production system at <a href="https://zuericitygpt.ch">zuericitygpt.ch</a> switched from GPT-4o-mini to GPT-5.6 Luna — the cost-efficient tier in OpenAI's latest model family. So in the following, we will take a look at the effect of two simultaneous model upgrades on ZueriCityGPT, tested against the same knowledge base.</p>
<h2>What is new in each model</h2>
<p><strong>Apertus 1.5</strong> adds 2 trillion tokens of continued pretraining to the 70B model. For the first time, the model accepts images alongside text, and audio input is available experimentally. Instruction following — a clear weak spot in 1.0 — was a focus of the update. The model remains fully open-source under Apache 2.0, trained on the Alps supercomputer at CSCS.</p>
<p><strong>GPT-5.6 Luna</strong> is the smallest model in OpenAI's <a href="https://openai.com/index/gpt-5-6/">GPT-5.6 family</a> (Sol / Terra / Luna), released on July 9, 2026. It supports a million-token context window and is designed for high-volume, latency-sensitive workloads — the profile of a public-sector chatbot like ZüriCityGPT. The production system ran on GPT-5.4-nano for most of July before switching to Luna on July 31; our controlled test on August 7 ran against Luna.</p>
<h2>A disclaimer about the comparison</h2>
<p>An important point needs to be reiterated, as it was often misunderstood in the first Apertus release: it is not possible to directly compare the performance of a Large Language Model (LLM) like Apertus to an integrated solution like Luna, which does not provide access directly to the closed LLM inside of it. The software stack that is run by OpenAI to optimize interactions with their models is proprietary, whereas Apertus is running on an open-source stack that you can deploy in the cloud or on-premises.</p>
<p>Similarly, while ZueriCityGPT has an OSS model version, not all source code of the underlying system LiipGPT has been open-sourced at this point. LiipGPT ships with a lot of features that a sophisticated Retrieval-Augmented-Generation (RAG)-system needs in order to provide you with accurate answers based on a curated knowledge base.</p>
<p>Both models run on the same <a href="https://liipgpt.ch">LiipGPT</a> platform and the same knowledge base of official content from <a href="https://www.stadt-zuerich.ch">stadt-zuerich.ch</a>. Note that the OSS version also uses different models for embedding (Qwen3-Embedding-0.6B) and reranking (BGE-Reranker), which can affect answer quality. Neither system uses web search — all answers are generated exclusively from documents retrieved from the stadt-zuerich.ch knowledge base via RAG.</p>
<p>The OSS version of ZüriCityGPT will add 16k tokens from the chunks into the prompt while the Luna version adds 32k tokens. This is because the price of Apertus tokens is roughly 4 times higher than Luna model tokens. We therefore actively choose to constrain the amount of tokens used for cost-control and to achieve a similar cost to question ratio.</p>
<h2>The numbers</h2>
<p>We have two data sources: historical quality metrics from real user traffic (measured automatically by the LiipGPT platform), and a controlled test where we asked 20 identical questions to both systems on the same day.</p>
<p>A note on how to read the comparisons: throughout this post, we make two kinds of comparison. <strong>Before → after</strong> compares the same system across its model upgrade (Apertus 1.0 → 1.5, or GPT-4o-mini → GPT-5.6 Luna). <strong>Head-to-head</strong> compares the two systems against each other on the same question, same day. We always name the specific models so it is clear which comparison we are making.</p>
<h3>Historical metrics — real user traffic</h3>
<p>Over the past months, the LiipGPT platform has scored every answer using our three established metrics:</p>
<ul>
<li>Match Rate (MR: is the answer acceptable?)</li>
<li>Match Rate +1 (MR+1: is the answer actually good?)</li>
<li>Faithfulness (Faith: does the answer stick to the source documents?).</li>
</ul>
<p>In the following, we aggregate each model's entire deployment period. This gives the clearest before-and-after picture.</p>
<figure><a href="https://liip.rokka.io/dynamic/b60ce6/chart-comparison-fullrange.png"><img alt="" src="https://liip.rokka.io/www_inarticle_5/b60ce6/chart-comparison-fullrange.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/b60ce6/chart-comparison-fullrange.jpg 2x"></a><figcaption>Full-range model comparison from real user traffic.</figcaption></figure>
<table>
<thead>
<tr>
<th>Period</th>
<th>System</th>
<th>Model</th>
<th>n</th>
<th>Avg Score</th>
<th>MR</th>
<th>MR+1</th>
<th>Faith</th>
</tr>
</thead>
<tbody>
<tr>
<td>Jan–Jun 2026</td>
<td>PROD</td>
<td>GPT-4o-mini</td>
<td>3,935</td>
<td>4.44</td>
<td>98.7%</td>
<td>70.3%</td>
<td>0.83</td>
</tr>
<tr>
<td>Jul 31–Aug 14</td>
<td>PROD</td>
<td>GPT-5.6 Luna</td>
<td>216</td>
<td>4.81</td>
<td>99.1%</td>
<td>86.6%</td>
<td>0.84</td>
</tr>
<tr>
<td>Jan–Jun 2026</td>
<td>OSS</td>
<td>Apertus 1.0</td>
<td>1,466</td>
<td>3.73</td>
<td>91.5%</td>
<td>40.2%</td>
<td>0.53</td>
</tr>
<tr>
<td>Jul 25–Aug 14</td>
<td>OSS</td>
<td>Apertus 1.5</td>
<td>287</td>
<td>3.99</td>
<td>87.1%</td>
<td>59.2%</td>
<td>0.56</td>
</tr>
</tbody>
</table>
<p>Apertus 1.5 made the bigger relative leap: Match Rate +1 jumped from 40% to 59% — a 19 percentage point improvement. GPT-5.6 Luna went from 70% to 87%, a 16 percentage point jump. The absolute gap narrowed only slightly, from 30 points to 27.</p>
<h3>Controlled test — 20 identical questions</h3>
<p>We repeated the exact same <a href="https://www.liip.ch/en/blog/apertus-after-8-months-what-we-learned-and-look-forward-using-switzerlands-ai-model#the-20-question-test">20 questions from the original blog post</a>, asking them on both systems on August 7, 2026. One person asked all 20 questions on each system — this is not a user study but a controlled comparison isolating model changes from knowledge-base changes.</p>
<figure><a href="https://liip.rokka.io/dynamic/5ee10f/chart-comparison-controlled.png"><img alt="" src="https://liip.rokka.io/www_inarticle_5/5ee10f/chart-comparison-controlled.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/5ee10f/chart-comparison-controlled.jpg 2x"></a><figcaption>Three quality metrics compared across the controlled test — 20 identical questions.</figcaption></figure>
<table>
<thead>
<tr>
<th>#</th>
<th>Question</th>
<th>Apertus 1.0</th>
<th>Apertus 1.5</th>
<th>GPT-4o-mini</th>
<th>GPT-5.6 Luna</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>Who is the mayor of the city?</td>
<td>3</td>
<td>2</td>
<td>5</td>
<td>5</td>
</tr>
<tr>
<td>2</td>
<td>which day will paper get collected in 8004</td>
<td>6</td>
<td>6</td>
<td>5</td>
<td>6</td>
</tr>
<tr>
<td>3</td>
<td>do you understand english?</td>
<td>6</td>
<td>6</td>
<td>6</td>
<td>6</td>
</tr>
<tr>
<td>4</td>
<td>was gibts neues im zoo?</td>
<td>3</td>
<td>5</td>
<td>5</td>
<td>5</td>
</tr>
<tr>
<td>5</td>
<td>what are the stadtammanns doing?</td>
<td>5</td>
<td>5</td>
<td>5</td>
<td>5</td>
</tr>
<tr>
<td>6</td>
<td>what do you know?</td>
<td>3</td>
<td>3</td>
<td>5</td>
<td>4</td>
</tr>
<tr>
<td>7</td>
<td>What are the rules to use the Zurich waste dump?</td>
<td>6</td>
<td>3</td>
<td>3</td>
<td>6</td>
</tr>
<tr>
<td>8</td>
<td>wie alt ist die stadt zürich?</td>
<td>5</td>
<td>6</td>
<td>5</td>
<td>5</td>
</tr>
<tr>
<td>9</td>
<td>can you speak english?</td>
<td>3</td>
<td>5</td>
<td>3</td>
<td>5</td>
</tr>
<tr>
<td>10</td>
<td>Wer bist du?</td>
<td>3</td>
<td>6</td>
<td>6</td>
<td>1</td>
</tr>
<tr>
<td>11</td>
<td>Est-ce que tu parles français?</td>
<td>3</td>
<td>6</td>
<td>5</td>
<td>6</td>
</tr>
<tr>
<td>12</td>
<td>Was macht Corine Mauch?</td>
<td>5</td>
<td>5</td>
<td>4</td>
<td>5</td>
</tr>
<tr>
<td>13</td>
<td>What is Smart City?</td>
<td>4</td>
<td>5</td>
<td>4</td>
<td>5</td>
</tr>
<tr>
<td>14</td>
<td>when is the next zürifest</td>
<td>5</td>
<td>5</td>
<td>5</td>
<td>5</td>
</tr>
<tr>
<td>15</td>
<td>How many people live in Zurich</td>
<td>5</td>
<td>6</td>
<td>5</td>
<td>6</td>
</tr>
<tr>
<td>16</td>
<td>How do I dispose of Teflon pans?</td>
<td>3</td>
<td>2</td>
<td>4</td>
<td>5</td>
</tr>
<tr>
<td>17</td>
<td>What is Zurich doing for climate protection?</td>
<td>5</td>
<td>5</td>
<td>6</td>
<td>6</td>
</tr>
<tr>
<td>18</td>
<td>Where do I do my tax declaration?</td>
<td>4</td>
<td>5</td>
<td>5</td>
<td>5</td>
</tr>
<tr>
<td>19</td>
<td>Who is the head of the OIZ?</td>
<td>6</td>
<td>6</td>
<td>6</td>
<td>6</td>
</tr>
<tr>
<td>20</td>
<td>How old is the city of Zurich?</td>
<td>5</td>
<td>2</td>
<td>6</td>
<td>6</td>
</tr>
<tr>
<td></td>
<td><strong>Average</strong></td>
<td><strong>4.40</strong></td>
<td><strong>4.70</strong></td>
<td><strong>4.90</strong></td>
<td><strong>5.15</strong></td>
</tr>
</tbody>
</table>
<p>Apertus 1.5 improved its average score from 4.40 to 4.70. GPT-5.6 Luna came in at 5.15.</p>
<h4>How did the three key metrics change?</h4>
<p>Our automated scoring scale runs from 1 (wrong or refuses to answer) through 3 (partially correct) to 6 (excellent). We define "acceptable" as any score above 2 — the answer is at least partially useful. "Good" means above 3 — the answer is substantively correct and helpful, not just borderline.</p>
<table>
<thead>
<tr>
<th>Metric</th>
<th>Apertus 1.0</th>
<th>Apertus 1.5</th>
<th>Δ</th>
<th>GPT-4o-mini</th>
<th>GPT-5.6 Luna</th>
<th>Δ</th>
</tr>
</thead>
<tbody>
<tr>
<td>Acceptable (&gt;2)</td>
<td>100% (20/20)</td>
<td>85% (17/20)</td>
<td>−15pp</td>
<td>100% (20/20)</td>
<td>95% (19/20)</td>
<td>−5pp</td>
</tr>
<tr>
<td>Good (&gt;3)</td>
<td>65% (13/20)</td>
<td>75% (15/20)</td>
<td>+10pp</td>
<td>90% (18/20)</td>
<td>95% (19/20)</td>
<td>+5pp</td>
</tr>
<tr>
<td>Faithfulness</td>
<td>0.60</td>
<td>0.61</td>
<td>+0.01</td>
<td>0.85</td>
<td>0.80</td>
<td>−0.05*</td>
</tr>
</tbody>
</table>
<p>*The faithfulness delta for Luna is a measurement artifact — see the discussion in "What we learned" below.</p>
<p>Apertus 1.5 gives fewer acceptable answers (85% vs 100%) but more good ones (75% vs 65%). This is the "honest refusal" pattern: the model no longer produces borderline answers that scrape past the threshold. When it answers, it answers well; when it cannot, it says so. The Apertus team confirmed this matches their own testing with OR-Bench, a benchmark designed to measure over-refusal behaviour. GPT-5.6 Luna shows a milder version of the same trend.</p>
<p>In a head-to-head battle, GPT-5.6 Luna won 6 questions, Apertus 1.5 won 2, and 12 were ties. Compare that to the original: GPT-4o-mini won 9, Apertus 1.0 won 3, with 8 ties. Ties grew from 8 to 12 — meaning Apertus now matches the production model on most questions.</p>
<h3>Matched questions — real users, same question, both systems</h3>
<p>Twenty-five real user questions were asked on both systems during the Apertus 1.5 period. The questions from different users who but we chose to simulate them on both systems to achieve a comparison.</p>
<p>On matched questions, the gap in average score is just 0.16 points (4.92 vs 5.08). Match Rate: 96% vs 100%. Match Rate +1: 76% vs 92%. Faithfulness: 0.59 vs 0.83.</p>
<h2>The example questions revisited</h2>
<p>In the original blog, we highlighted specific scenarios to make the numbers tangible. Here is how they changed — plus two new ones that show the practical impact of both upgrades.</p>
<h3>"Est-ce que tu parles français?" — both answer now</h3>
<p>This was the most visible failure in the original comparison. When asked "Do you speak French?", GPT-4o-mini responded fluently in French. Apertus 1.0 declined, saying it could not answer questions about its own linguistic capabilities.</p>
<p>Apertus 1.5 now answers: <em>"Oui, je parle français. Je suis à votre disposition pour vous aider dans la langue que vous préférez."</em> The IFEval improvement that the Apertus team targeted has real practical impact.</p>
<figure><a href="https://liip.rokka.io/dynamic/daed48d/screenshot-french-q11.png"><img alt="" src="https://liip.rokka.io/www_inarticle_5/daed48d/screenshot-french-q11.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/daed48d/screenshot-french-q11.jpg 2x"></a><figcaption>Both systems now respond fluently in French — GPT-5.6 Luna left and Apertus 1.5 right.</figcaption></figure>
<h3>Population — improved but still off</h3>
<p>In the original test, GPT-4o-mini returned Zurich's current population (452,421), while Apertus gave a 2013 figure. Apertus 1.5 now returns a more up-to-date number but uses the wrong date. GPT-5.6 Luna cites the correct combination of time and population. Apertus 1.5 plays a part in this, but retrieval logs tell a different story: the OSS embedding model (Qwen3-Embedding-0.6B) ranks old statistical PDFs from 2010–2013 above current population pages, and because those PDFs are large, they fill the smaller context window (16k vs production) before up-to-date sources are included. The production system's embedding model (OpenAI text-embedding-3-small) retrieves the current data as its top result.</p>
<figure><a href="https://liip.rokka.io/dynamic/1a0dda/screenshot-population-q15.png"><img alt="" src="https://liip.rokka.io/www_inarticle_5/1a0dda/screenshot-population-q15.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/1a0dda/screenshot-population-q15.jpg 2x"></a><figcaption>Both models return the current population figure, but Apertus cites the wrong date — GPT-5.6 Luna left and Apertus 1.5 right.</figcaption></figure>
<h3>Waste dump — the tables turned</h3>
<p>This was Apertus 1.0's showcase win in the original blog: scoring 6/6 where GPT-4o-mini scored 3/6. In 1.5, the roles reversed. GPT-5.6 Luna now delivers a comprehensive answer. Apertus 1.5 asks the user to specify what kind of waste they mean. This is the clearest example of Apertus 1.5's new pattern: better instruction following sometimes means it asks clarifying questions rather than making assumptions.</p>
<figure><a href="https://liip.rokka.io/dynamic/d90d8e/screenshot-waste-q7.png"><img alt="" src="https://liip.rokka.io/www_inarticle_5/d90d8e/screenshot-waste-q7.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/d90d8e/screenshot-waste-q7.jpg 2x"></a><figcaption>GPT-5.6 Luna now provides detailed waste disposal information left. Apertus 1.5 asks for clarification right where 1.0 gave a direct answer.</figcaption></figure>
<h3>"Who is the mayor?" — the honest refusal in practice</h3>
<p>This question illustrates why acceptable answers dropped from 100% to 85%. GPT-5.6 Luna correctly answered "Raphael Golta". Apertus 1.5 refused: "The provided information does not contain the name of the current mayor of the city." Rather than guessing like 1.0 did (which returned the outdated name "Corine Mauch"), 1.5 admits the gap. This is a retrieval issue, likely amplified by the smaller OSS embedding model.</p>
<figure><a href="https://liip.rokka.io/dynamic/5006d9/screenshot-mayor-q1.png"><img alt="" src="https://liip.rokka.io/www_inarticle_5/5006d9/screenshot-mayor-q1.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/5006d9/screenshot-mayor-q1.jpg 2x"></a><figcaption>GPT-5.6 Luna correctly names Raphael Golta left. Apertus 1.5 refuses — it cannot find the name in the retrieved documents right.</figcaption></figure>
<h3>"was gibts neues im zoo?" — Apertus finds its voice in German</h3>
<p>Apertus 1.0 scored 3/6 — a vague, generic answer. Apertus 1.5 now delivers a detailed German response covering the Lewa Savanne, the Kaeng Krachan Elefantenpark, and the Zooseilbahn controversy. It scores 5/6, matching GPT-5.6 Luna.</p>
<p>This is notable because the knowledge base is predominantly German, and zoo news is the kind of practical, frequently-asked city question where a chatbot needs to perform. Apertus 1.0 struggled to synthesize multiple retrieved documents into a coherent answer; 1.5 does this naturally. The improvement is not about finding the right documents — the RAG pipeline retrieved them before, too — but about what the model does with them once retrieved.</p>
<figure><a href="https://liip.rokka.io/dynamic/963092/screenshot-zoo-q4.jpg"><img alt="" src="https://liip.rokka.io/www_inarticle_5/963092/screenshot-zoo-q4.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/963092/screenshot-zoo-q4.jpg 2x"></a><figcaption>Both systems now give detailed zoo news — GPT-5.6 Luna left, Apertus 1.5 right. Both score 5/6.</figcaption></figure>
<h2>What we learned</h2>
<p><strong>The quality gap is closing, from both sides.</strong> Apertus 1.5 narrowed the gap substantially — but both models got better at the same time. On matched real-user questions, the average score difference is just 0.16 points. On the controlled test, the gap went from 0.50 (4.40 vs 4.90) to 0.45 (4.70 vs 5.15).</p>
<p><strong>Instruction following improved — with a tradeoff.</strong> Apertus 1.0 had a clear weakness in instruction following (IFEval 44%). In 1.5, the French language question works, the identity answer is clean, and the English responses are more natural. But the same improvement created a new pattern: when retrieved documents do not clearly contain the answer, the model now refuses rather than attempting to synthesize. This dropped the controlled-test Match Rate from 100% to 85%.</p>
<p><strong>The English/German gap persists — but it is not the model.</strong> We asked "wie alt ist die stadt zürich?" in German and "How old is the city of Zurich?" in English. Apertus 1.5 scored 6/6 on the German version and 2/6 on the English version. The gap most likely comes from the OSS embedding model struggling with cross-lingual retrieval against the predominantly German knowledge base.</p>
<p><strong>Faithfulness requires careful measurement.</strong> The faithfulness numbers initially included refusal answers, which led to low scores because the refusal obviously cannot be grounded in facts. On that basis, GPT-5.6 Luna scores 0.84 — essentially identical to GPT-4o-mini's 0.83. Apertus improved slightly (0.53 to 0.56 in production traffic, 0.60 to 0.61 in the controlled test). "Faithfulness" is our platform's metric for how well the answer sticks to the retrieved sources — it is not a term the Apertus team uses internally, but they confirmed that general accuracy improvement and alignment work are ongoing goals. The gap between Apertus and the production system (0.56 vs 0.84) depends as much on how the RAG pipeline surfaces and frames source material as on the model itself.</p>
<p><strong>Both systems benefit equally from knowledge base updates.</strong> The knowledge base now correctly reflects that Raphael Golta replaced Corine Mauch as Stadtpräsident in May 2026. Both models handle this correctly — confirming that the LiipGPT platform's RAG pipeline works as designed: model-agnostic, with the knowledge base as the single source of truth.</p>
<h2>Looking ahead</h2>
<p>Compared to our recent test of Apertus 1.0, the picture has shifted meaningfully. Back then, Apertus was an experiment with many limitations. Apertus 1.5 seems to handle many questions at the same level as GPT-5.6 Luna. The cases where it still falls short (faithfulness, cross-lingual retrieval) are increasingly about other open-source components involved, less on the model itself. But in our test field of RAG application, the cost ratio is important. If you compare the token cost of GPT-5.6 Luna on Azure with the token cost of Apertus 1.5 70B on Infomaniak, the Apertus model will cost you roughly 4 times as much as the GPT model. Obviously there are other advantages like enhanced data sovereignty and higher ethical standards in training the model but still, being able to run Apertus 1.5 in sizes between 8B and 70B would help achieve a better cost to value ratio in our case.</p>
<p>We are sharing these results with the Apertus team and the wider community for review. As before, our aim is not to provide a scientific benchmark but a practical report from a production deployment. Consider that these are just our first impressions, and we are looking forward to further test Apertus with other RAG-deployments and in more use case scenarios. If you would like to test how Apertus 1.5 works when refining brand text, our latest <a href="https://liip-textmate.liipgpt.ch/">TextMate</a> has the model available as well.</p>
<p>The question from our <a href="https://www.liip.ch/en/blog/apertus-after-8-months-what-we-learned-and-look-forward-using-switzerlands-ai-model">last post</a> — "how much quality gap are you willing to accept for full sovereignty?" — now has a different answer. On most questions, the answer quality is equivalent. The remaining gap is in faithfulness and source attribution — Apertus is improving, and GPT-5.6 Luna maintains the same level as its predecessor once measurement artifacts are accounted for. Closing that gap will require improvements to the open-source RAG stack (embedding, reranking) as well as finding a good mix between model size and speed / price and might be less focused on the language model itself.</p>
<p>Please do not hesitate to <a href="https://www.liip.ch/en/team/josef-kruckenberg">contact me</a> if you have any questions about Apertus or AI solutions in general; I’d be delighted to discuss this with you. </p>
<p><strong>Acknowledgements</strong></p>
<p>Thank you to the <a href="https://ai.ethz.ch/news-and-events/ai-center-news/2026/07/apertus-15-building-the-next-generation-of-open-ai-infrastructure.html">Swiss AI Initiative</a> — EPFL, ETH Zurich, and CSCS — for continuing to develop Apertus as a public good. Thank you Oleg Lavrovsky and Martin Renou for reviewing a draft of this post and providing feedback from the Apertus team's perspective.</p>
<p>Thank you <a href="https://www.liip.ch/de/team/chregu">Christian Stocker</a> and the <a href="https://liipgpt.ch">LiipGPT</a> team for deploying the update within 24 hours of release.</p>
<p>Thank you <a href="https://publicai.ch/">Public AI</a> and <a href="https://www.infomaniak.com/en/hosting/ai-services/open-source-models">Infomaniak</a> for hosting Apertus inference for us.</p>
<p>Parts of this analysis were prepared with the help of Claude. The data was collected and scored automatically by the LiipGPT platform.</p>]]></description>
    </item>
        <item>
      <title>Von Circle zu Circle, ohne sich im Kreis zu drehen: mein erstes Jahr bei Liip</title>
      <link>https://www.liip.ch/de/blog/von-circle-zu-circle-ohne-sich-im-kreis-zu-drehen-mein-erstes-jahr-bei-liip</link>
      <guid>https://www.liip.ch/de/blog/von-circle-zu-circle-ohne-sich-im-kreis-zu-drehen-mein-erstes-jahr-bei-liip</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 +0200</pubDate>
      <description><![CDATA[<p>Das war kein Zufall. Das hat meine Berufsbildnerin Petra so geplant. Sie wollte, dass ich möglichst viele Circles und Liiper kennenlerne und dabei so viele Bildungsziele wie möglich abhake. Am Ende zählte, wo es gerade Arbeit und Betreuung für mich gab, nicht ein fixer Plan auf Papier. Trotzdem hat sich im Rückblick eine klare Reihenfolge ergeben, und genau die hat den Unterschied gemacht. Jeder Circle hat mir etwas mitgegeben, das ich im nächsten gebraucht habe.</p>
<h3>Vom 1. August bis 1. November war ich im Finance Circle. Hier hat alles angefangen</h3>
<p>Ich habe den Umzug des Finance Bereichs weg von unserem Atlassian Wiki begleitet. Die neue Lösung stand damals noch nicht fest. Darum habe ich erstmal alles davon auf Google Drive abgelegt. Das klingt nach einer einfachen Aufräumaktion. Tatsächlich war es ein guter Einstieg, um zu sehen, wie viel Wissen bei Liip überhaupt zusammenkommt und wo es abgelegt ist.</p>
<p>Meine Aufgaben im Circle:</p>
<ul>
<li>Verantwortung für unsere Finance E-Mail, bei der alle Rechnungen ankommen. Ich habe sie abgeklärt und ins RMA, unser Finance System, eingetragen.</li>
<li>Monatliche Kreditkartenabrechnung, bei der ich jedes Mal auch nach Automatisierungspotential gesucht habe.</li>
<li>Auffrischung der NDA Ablage. Ich habe die ganze Historie übernommen und in Google Sheets ein paar Automatisierungen dafür gebaut, damit niemand mehr von Hand nachschauen muss, wo welcher Vertrag liegt.</li>
</ul>
<p>Am meisten Spass hatte ich, wenn ich frei nach Automatisierungspotential suchen und dazu recherchieren durfte. Weniger spannend waren die grossen Datenübertragungen. Ein guter Tipp dazu: Die drei Fragezeichen laufen lassen, dann geht das auch. Die typischste Finance Aufgabe war für mich die monatliche Kreditkartenabrechnung. Genau die Art Arbeit, die ein Unternehmen im Hintergrund am Laufen hält.</p>
<p><strong>Im Finance Circle habe ich nicht in erster Linie Zahlen gelernt, sondern wie Liip als Unternehmen tickt.</strong></p>
<h3>Vom 1. November bis 1. Februar war ich im <a href="https://www.liip.ch/de/services/development/cms/drupal">Drupal</a> Zürich Circle. Hier wurde es konkreter</h3>
<p>Hier durfte ich zum ersten Mal ein Stück PO Rolle einnehmen. Zusammen mit Petra habe ich ein Projekt begleitet. Bei anderen Projekten habe ich zuerst zugeschaut und kleinere Aufgaben übernommen, um das Business dahinter zu verstehen. Das war ein anderes Tempo als im Finance Circle. Statt bestehende Prozesse zu übernehmen, musste ich verstehen, wie ein Kundenprojekt überhaupt entsteht, von der ersten Anforderung bis zur Umsetzung.</p>
<p>Später kamen grössere Aufgaben dazu:</p>
<ul>
<li>PO Testing nach der ersten Umsetzung durch die Entwickelnden</li>
<li>Ein Handbuch für das Drupal Backend, das ich von Grund auf geschrieben habe</li>
</ul>
<p>Bei diesem Handbuch mussten die Erklärungen so einfach sein, dass auch jemand ohne Coding Hintergrund damit arbeiten kann. Genau das war für mich die grösste Übung.</p>
<figure><img alt="" src="https://liip.rokka.io/www_inarticle_5/0b07ea/liipgpt-manual-de.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/0b07ea/liipgpt-manual-de.jpg 2x"></figure>
<p>Auch wenn das bei Liip ein Production Circle ist, habe ich dort nie selbst programmiert. In der Berufsschule hatte ich zwar mit Python zu tun. Das war aber nicht wirklich meins. Was ich sehr schätze: Petra pusht mich in unseren Syncs immer wieder Richtung Code und erklärt oder zeigt mir Dinge. Im Circle selbst hatte ich viel mit Backend und Frontend zu tun, aber nie im Code direkt. Am coolsten war für mich der Moment, als ich das Handbuch fertiggestellt hatte. Danach kamen viele positive, dankbare Reaktionen.</p>
<p><strong>Im Drupal Zürich Circle habe ich gelernt, wofür Liip als Firma eigentlich steht.</strong></p>
<h3>Vom 1. Februar bis 1. August war ich im Design und Content Circle. Hier durfte ich zeigen, was ich gelernt hatte</h3>
<p>In diesem Circle drehte sich vieles um KI mit Fokus auf Content, sei es Erstellung, Aufbereitung oder Guidelines. Für Kundenprojekte durfte ich:</p>
<ul>
<li>Promptsheets für <a href="https://liip-textmate.liipgpt.ch/">Textmates</a> erfassen</li>
<li>Prozessoptimierungen ausprobieren</li>
<li><a href="https://www.liip.ch/de/blog/who-still-needs-an-apprenticeship-when-ai-exists">Einen eigenen Blogpost schreiben</a></li>
<li>Eigene KI Tools bauen</li>
<li>Bei Schulungen und Workshops mithelfen und sie am Ende selbst mitgestalten</li>
</ul>
<p>Das reicht von der Vorbereitung der Unterlagen bis zur aktiven Unterstützung während des Workshops selbst, wenn Teilnehmende Fragen haben oder an einer Übung feststecken.</p>
<p>Der Workshop, den ich eingangs erwähnt habe, war mit Caritas St. Gallen und Appenzell. Gleichzeitig war es unser allererster Responsible AI Workshop. Genau in dem Jahr, in dem ich selbst bei Liip gewachsen bin, hat sich auch unsere Expertise zu diesem Thema stark weiterentwickelt.</p>
<figure><img alt="" src="https://liip.rokka.io/www_inarticle_5/5a4eb5/responsible-ai-workshop.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/5a4eb5/responsible-ai-workshop.jpg 2x"></figure>
<p>Ich war so lange in keinem anderen Circle. Genau das hat einen Unterschied gemacht. Ich war viel stärker in die Arbeit des Circles integriert. Vor allem in der zweiten Hälfte habe ich gemerkt, wie viel Vertrauen und Verantwortung ich bekommen habe, um Dinge alleine umzusetzen. Inhaltlich war die Arbeit hier besonders spannend: Es ging viel um die interne Gestaltung von Liip, kombiniert mit Workshops für Kunden, bei denen wir unser Wissen weitergeben konnten. Das war nochmal etwas anderes als in den beiden Circles davor.</p>
<p><strong>Im Design und Content Circle durfte ich nicht mehr nur lernen, sondern selbst etwas weitergeben.</strong></p>
<h3>Was sich über das Jahr aufgebaut hat</h3>
<p>Wenn ich die drei Circles nebeneinander lege, sehe ich vor allem eine Kurve bei der Verantwortung. Im Finance Circle habe ich Prozesse übernommen, die schon standen, und dabei Leute, Tools und Abläufe bei Liip kennengelernt. Im Drupal Zürich Circle ging es dann mehr um das eigentliche Geschäft von Liip, die Webentwicklung, und ich durfte bei PO Testing schon eigene Entscheidungen treffen und sogar Kunden kennenlernen. Im Design und Content Circle konnte ich dieses Wissen nutzen, um bei strategischen Aufgaben mitzuarbeiten, eigene Tools zu bauen und Workshops mitzugestalten.</p>
<p>Jeder Circle hat mich auf den nächsten vorbereitet. Nicht, weil ich das selbst geplant hätte, sondern weil genau das System hinter der Ausbildung zum Entwickler Digitales Business bei Liip ist. Am deutlichsten sehe ich das bei der Automatisierung. Im Finance Circle habe ich bei der Kreditkartenabrechnung gelernt, wie man einen Prozess auseinandernimmt und nach Automatisierungspotenzial sucht. Genau dieses Denken habe ich im Design und Content Circle wieder gebraucht, als ich eigene KI Tools gebaut habe. Ohne den ersten Schritt hätte ich den dritten nicht so angehen können.</p>
<p>Eine Sache hat sich dabei durch alle drei Circles gezogen, und die hat nichts mit Fachwissen zu tun. Ich habe gelernt, dass ich auf andere zugehen und um Hilfe fragen kann. Nicht, weil ich Lernender bin und das erlaubt ist, sondern weil das bei Liip für alle selbstverständlich ist.</p>
<p><strong>Vertrauen kam bei mir nicht auf einmal, sondern Circle für Circle dazu, und genau das hat mich vom letzten August bis heute am meisten verändert.</strong></p>
<p>Wie sieht deine eigene Lernkurve aus? Baut sich das bei dir auch Stück für Stück auf, oder eher sprunghaft?</p>]]></description>
    </item>
        <item>
      <title>Deine Content Strategie als System-Prompt</title>
      <link>https://www.liip.ch/de/blog/deine-content-strategie-als-system-prompt</link>
      <guid>https://www.liip.ch/de/blog/deine-content-strategie-als-system-prompt</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 +0200</pubDate>
      <description><![CDATA[<p>Die Kosten für Content-Produktion tendieren heute gegen Null. </p>
<p>Doch wer im Bereich Branding- und Unternehmenkommunikation arbeitet kennt das Dilemma: </p>
<ul>
<li>Schneller und konsistenter: Ja gerne!</li>
<li>Kontrolle über die Positionierung verlieren: Nein danke.</li>
</ul>
<p>Wir haben unsere eigene Content Strategie überarbeitet – mit der Hilfe von KI. Und einen Weg gesucht, das Dilemma aufzulösen: Ein effizienter Prozess, der die Kontrolle beim Menschen belässt.</p>
<h1>Womit starten?</h1>
<p>Liip hat seine Positionierungsziele geschärft, worauf wir auch unsere Content Strategie überarbeiten mussten. </p>
<p>Für die Überarbeitung habe ich mit Claude gearbeitet: Wir haben einen Enterprise-Zugang zu Claude, ich arbeite mit der Desktop App. </p>
<p>In einem ersten Schritt habe ich ein neues Projekt angelegt. Ich gebe dem Projekt einen prägenden Titel und beschreibe, was ich vorhabe. </p>
<p>Ich kann dem System ausserdem <strong>Kontext-Informationen</strong> geben, die es bei allen Arbeiten berücksichtig.</p>
<p>In diesem Fall relevant: </p>
<ul>
<li>Vision und Strategie der Organisation</li>
<li>Positionierungsziele</li>
<li>Bestehende Personas</li>
<li>Bestehende Freigabeprozesse und Stakeholder, die am Ende zustimmen müssen</li>
<li>...</li>
</ul>
<p>Alles, was für die Content Strategie relevant ist, kommt als Kontext ins Projekt. Gerade wenn mehrere Stakeholder oder Sprachregionen am Ende ein Wort mitzureden haben, macht dieser Aspekt den Unterschied: Das System schlägt vor – aber es kennt die Leitplanken, innerhalb derer ein Vorschlag überhaupt tragfähig ist.</p>
<p>Diese Informationen bieten die solide Grundlage für ein <strong>verankertes</strong> und <strong>effizientes</strong> Vorgehen. </p>
<p>Nun kann die Arbeit an der eigentlichen <a href="https://www.liip.ch/de/blog/das-gehoert-in-deine-content-strategie" rel="noreferrer" target="_blank">Content Strategie</a> starten. </p>
<p>Ich wollte unsere Kernbotschaft und Zielgruppen schärfen – zwei Bausteine, die wir für die unsere Positionierung weiterentwickelt hatten.</p>
<h1>Zusammenarbeit mit dem System?</h1>
<p>Bei der Überarbeitung jedes Bausteins hatte ich jeweils zwei Optionen:</p>
<ul>
<li>Entweder lasse ich mir von der KI einen ersten Vorschlag generieren, den das System aus den Kontextinformationen ableitet. </li>
<li>Oder ich skizzierte grob, was mir vorschwebt, und das System generiert eine vollständige Version.</li>
</ul>
<p>Beide Ansätze geben mir einen Startpunkt, den ich anschliessend weiter verfeinern kann – allein, im Tandem mit dem System oder mit meinem Team. Die KI schreibt also nicht meine Strategie, sondern unterstützt mich im Prozess. </p>
<p><strong>Es braucht dabei den Moment, in dem ich von den generierten Ergebnissen zurücktrete. Reflektiere, ob diese meinen Zielen wirklich gerecht werden. Und meine eigenen Prioritäten einbringe.</strong> </p>
<p>Um einen frischen Blick auf generierte Vorschläge zu erhalten, hole ich mir gerne Mitarbeitende hinzu. Ich habe in diesem Prozess eine erste vollständige Version unserer Kernbotschaften entwickelt, zu der ich mir anschliessend Feedback aus dem Team hole. </p>
<p>Je nach Anspruch und Stand kannst du dich so durch die verschiedenen Themen arbeiten. Das System stellt sicher, dass die Inhalte sinnvoll aufeinander aufbauen. </p>
<h1>Was ist ein System-Prompt?</h1>
<p>Sobald meine überarbeitete Strategie steht, gehe ich noch einen Schritt weiter: Ich leite daraus einen System-Prompt ab.</p>
<p>Ein System-Prompt ist die übergeordnete Anweisung, die du einem Modell einmal mitgibst und die dann für alle folgenden Interaktionen in diesem Kontext automatisch gilt – im Unterschied zum eigentlichen Prompt, den du für jede einzelne Anfrage neu schreibst.</p>
<p>Damit habe ich mir im Grunde einen Assistenten gebaut, der mich – oder andere aus dem Team – bei neuen Inhalten unterstützt und dabei automatisch auf unsere Content Strategie einzahlt:</p>
<ul>
<li>Das System fragt mich, für welche Zielgruppe der Content gedacht ist, und liefert dazu passende Pain Points und Content-Bedürfnisse.</li>
<li>Das System fragt mich, welche Kernbotschaft im Vordergrund stehen soll und hilft mir, den Content daran auszurichten.</li>
<li>Das System führt mich durch alle weiteren relevanten Dimensionen – Schritt für Schritt.</li>
</ul>
<p>In der Arbeit mit dem System kann der System-Prompt dann weiter verfeinert werden (zum Beispiel, weil ich realisiere, dass es darauf ankommt, für welchen Kanal ich Content verfasse.)</p>
<h1>Mehr als ein Tool-Trick?</h1>
<p>Genau hier zeigt sich für mich der Punkt, der oft übersehen wird: <strong>Automatisierung skaliert nur, was strukturell schon stimmt.</strong> </p>
<p>Der System-Prompt funktioniert nur so gut wie die Content Strategie, aus der er abgeleitet ist. Die eigentliche Arbeit war nicht das Prompten, es war das Schärfen der Strategie. Die strategischen Entscheidungen – wofür wir stehen, wen wir erreichen wollen, mit welcher Botschaft – haben wir im Team getroffen.</p>
<p>Wenn du magst, probiere es an einem einzelnen Baustein aus: Kontext rein, ein Thema auswählen, Vorschlag oder Skizze durchgehen lassen, Finetuning mit einer Mitarbeiterin.</p>
<p>Wenn du bereits mit uns arbeitest und deine Content Strategie einen Reality-Check braucht, sprich uns an: Gemeinsam sammeln und sichten wir das vorhandene Material und setzen es dann in eurem KI-Tool nach Wahl auf.</p>
<p>Arbeitest du im öffentlichen Sektor und fragst dich, wie sich ein solcher Prozess mit euren Freigabe- und Compliance-Anforderungen vereinbaren lässt? Lass uns dazu austauschen, wir haben Erfahrung damit, KI-gestützte Prozesse innerhalb bestehender Governance-Strukturen aufzusetzen.</p>]]></description>
    </item>
        <item>
      <title>Reimagining Frontend Frameworks</title>
      <link>https://www.liip.ch/de/blog/reimagining-reactive-frontend-frameworks</link>
      <guid>https://www.liip.ch/de/blog/reimagining-reactive-frontend-frameworks</guid>
      <pubDate>Thu, 16 Jul 2026 00:00:00 +0200</pubDate>
      <description><![CDATA[<p>Lately React, Vue and Svelte have converged on similar concepts for building reactive interfaces. Their syntax differs, but every reactive frontend still uses some form of state, computed values, mutations and effects.</p>
<p>Nevertheless, there is a lot of knowledge involved in getting something basic working in any of these frameworks. In this blog post I will explore what the most stripped down version of such a framework could look like, while still keeping the ergonomics of a reactive framework.</p>
<p>Let's get right into it.</p>
<h1>Imagining the framework</h1>
<p>To make this concrete, I built a small framework called <a href="https://www.npmjs.com/package/effectlayer">effectlayer</a>. I will explain its building blocks in this blog post.</p>
<h1>Building blocks</h1>
<p>The framework is built around a single JavaScript class, which the framework enhances with reactive behavior. That reactivity is powered by Signals, a common primitive used in frameworks such as Vue, Svelte, and SolidJS.</p>
<h2>State</h2>
<p>All class properties are treated as state. The framework turns them into Signals, so no state annotations are needed here:</p>
<pre><code class="language-javascript">class MoodSwing {
  energy = 5;
  coffee = 0;
}</code></pre>
<h2>Computed Values</h2>
<p>Computed values use standard JavaScript getters. The framework treats them as derived Signals that update automatically when their dependencies change:</p>
<pre><code class="language-javascript">  get mood() {
    if (this.energy &gt; 8) return "🤪";
    if (this.energy &gt; 4) return "😀";
    if (this.energy &gt; 0) return "😑";
    return "😴";
  }</code></pre>
<h2>Mutations</h2>
<p>Methods are a natural fit for mutations:</p>
<pre><code class="language-javascript">  drinkCoffee() {
    this.coffee++;
    this.energy = Math.min(10, this.energy + 3);
  }

  work() {
    this.energy = Math.max(0, this.energy - 2);
  }</code></pre>
<h2>Effects</h2>
<p>The last concept we need is effects. Effects are a fancy way of saying "a thing that executes when dependent Signals change".</p>
<p>I chose methods starting with <code>$</code> for annotating them:</p>
<pre><code class="language-javascript">  $monitor() {
    if (this.coffee &gt; 10) console.warn("You may want to slow down.");
  }</code></pre>
<p>The framework sees that <code>$monitor()</code> uses <code>coffee</code>. After <code>coffee</code> changes, it calls the method again.</p>
<h2>HTML</h2>
<p>Rendering HTML is also an effect that just returns JSX.</p>
<pre><code class="language-javascript">  $ui() {
    return (
      &lt;main&gt;
        &lt;h1&gt;{this.mood}&lt;/h1&gt;
        &lt;button onClick={() =&gt; this.drinkCoffee()}&gt;☕ Coffee&lt;/button&gt;
        &lt;button onClick={() =&gt; this.work()}&gt;💻 Work&lt;/button&gt;
      &lt;/main&gt;
    );
  }</code></pre>
<p>All that is left is to make the class reactive by wrapping it in an <code>effectlayer</code> call:</p>
<pre><code class="language-javascript">const moodSwing = effectlayer(MoodSwing);
moodSwing.$monitor();
document.body.appendChild(moodSwing.$ui());</code></pre>
<p>Calling an effect once activates it. This means <code>$monitor()</code> will now run whenever <code>coffee</code> changes.</p>
<p>The <code>$ui()</code> call returns an HTML element and keeps it up to date.</p>
<h1>All the concepts we need</h1>
<p>So basically the framework only needs four concepts:</p>
<ul>
<li>properties for state</li>
<li>getters for computed values</li>
<li>methods for mutations</li>
<li>methods starting with <code>$</code> for effects</li>
</ul>
<p>If you want to try out this experimental framework use:</p>
<pre><code class="language-sh">npm create effectlayer</code></pre>
<p>Or check it out on <a href="https://www.npmjs.com/package/effectlayer">npm</a>.</p>
<h1>Deep Dive?</h1>
<p>Let me know if you want to see a deep dive into how I built this framework: <a href="mailto:falk.zwimpfer@liip.ch">falk.zwimpfer@liip.ch</a></p>]]></description>
    </item>
        <item>
      <title>Du willst KI? Darum brauchst du eine Content Strategie</title>
      <link>https://www.liip.ch/de/blog/du-willst-ki-darum-brauchst-du-eine-content-strategie</link>
      <guid>https://www.liip.ch/de/blog/du-willst-ki-darum-brauchst-du-eine-content-strategie</guid>
      <pubDate>Tue, 14 Jul 2026 00:00:00 +0200</pubDate>
      <description><![CDATA[<p>Um ehrlich zu sein: der Gen-AI-Hype bedroht mein Content Business. Wir sehen, dass Unternehmen Investitionen in ihren Content zurückhalten. Wieso? </p>
<p>Die Fähigkeiten Generativer KI machen klar, dass sich unser Berufsfeld stark verändern wird. Dieses ist jedoch breit: Verschiedene Kompetenzprofile und Tätigkeitsfelder spielen zusammen.</p>
<p>Somit ist die interessante Frage: Wie genau wird sich die Content-Arbeit in Organisationen verändern?</p>
<h1>Wandel der Profile</h1>
<p>Die Studie der AG CommTech und der GK Personalberatung befasst sich mit dem Wandel von Rollen und Kompetenzen in der Kommunikationsprofession. Ihre <a href="https://agcommtech.de/wp-portfolio/studie-kuenftige-rollen-und-kompetenzen-in-der-kommunikationsprofession-im-wandel-der-digitalisierung/" rel="noreferrer" target="_blank">Metastudie</a> zeigt: Gefragt ist heute eine Symbiose aus technologischer Effizienz und menschlicher Authentizität. Sie beschreiben einen "wachsenden Bedarf an Experten, die in der Lage sind, Daten zu interpretieren, daraus mit Hilfe von KI relevante Erkenntnisse abzuleiten und diese in effektive Kommunikationsstrategien zu übersetzen".</p>
<p><strong>Strategische Kompetenzen werden also noch wichtiger.</strong></p>
<h1>Quantität rauf, Qualität runter?</h1>
<p>Warum ist das so?<br />
Schauen wir es uns in einem sichtbaren Bereich an: generiertem Text.</p>
<p>Eine <a href="https://ai-on-the-internet.github.io/" rel="noreferrer" target="_blank">Standford-Studie</a> von Anfang 2026 zeigt, dass bereits rund 35% aller neu ins Netz gestellten Websites KI-generiert oder zumindest KI-assistiert entstanden sind. </p>
<p>Die Menge an generierten Inhalten nimmt schnell zu. Wie steht es um die Qualität?</p>
<p>Das ist keine simple Frage, da Qualität unterschiedliche Dimensionen hat. </p>
<p>Fragen wir mal Claude dazu. Der Bot findet eine <a href="https://contentmarketinginstitute.com/b2b-research/b2b-content-marketing-trends-research-2025#AI%20use%20trends" rel="noreferrer" target="_blank">Befragung des Content Marketing Instituts</a>, in der  nur 17% der B2B-Marketer die Qualität von KI-generiertem Content als exzellent oder sehr gut bewerten. 44% stufen sie als gut, 35% als "fair" und 4% als schlecht ein. Gleichzeitig vertrauen 67% dem KI-Output nur mittelmässig und 28% wenig. </p>
<p>Das ist zwar kein belastbarer Befund, deckt sich aber mit meiner Erfahrung. Lass uns darum der Frage nachgehen: Wie entsteht dieses mittelmässige Zeugnis?</p>
<h1>Input Qualität ≙ Output Qualität</h1>
<p>KI kann sehr vieles, aber sie kann nicht zaubern. <strong>LLMs brauchen eine solide Basis, um gute Ergebnisse zu generieren.</strong></p>
<p>Wenn ich für die Kanäle meines Unternehmens Content generieren will, bedeutet eine solide Basis Folgendes:</p>
<ul>
<li><strong>Vollständige Daten</strong> (die KI kann keinen verlässlichen Content generieren, wenn sie in deinem bestehenden Content dazu keine Informationen findet)</li>
<li><strong>Bedürfnisse der Zielgruppen und ihre Fragen</strong> sind zentraler Orientierungspunkt im Content</li>
<li><strong>Ein Content-Regelwerk</strong> personalisiert den Output und richtet ihn an den Kommunikationsregeln und -zielen der Organisation aus</li>
</ul>
<p>Das ist durchaus ernüchternd, wenn man erwartet hat, dass einem Gen-AI die Arbeit abnimmt. Denn je nach Ausgangslage generiert dieser Anspruch zuerst einmal neue Arbeit.</p>
<p><strong>Es zeigt aber auch, wo du ansetzen kannst, um die Qualität des generierten Contents zu steigern.</strong></p>
<p>Diese Anforderungen rücken klar <strong>Nutzerzentrierung</strong> und <strong>Content Strategie &amp; Guidelines</strong> in den Mittelpunkt – womit wir wieder bei den strategischen Kompetenzen sind.</p>
<p>In einem nächsten Beitrag werde ich beschreiben, was für mich in eine Content Strategie gehört, und welche Elemente eine Priorisierung verdienen.</p>]]></description>
    </item>
        <item>
      <title>Das geh&#246;rt in deine Content Strategie</title>
      <link>https://www.liip.ch/de/blog/das-gehoert-in-deine-content-strategie</link>
      <guid>https://www.liip.ch/de/blog/das-gehoert-in-deine-content-strategie</guid>
      <pubDate>Tue, 14 Jul 2026 00:00:00 +0200</pubDate>
      <description><![CDATA[<p>Generative KI bedeutet eine massive Disruption für das Content Business. Der gesamte Produktionsprozess wird durchgeschüttelt. Profile, Arbeitsprozesse, Tools – alles wird neu definiert.</p>
<p>Wenn die Kosten für Content-Produktion gegen Null tendieren, gewinnen zwei Punkte massiv an Bedeutung:</p>
<ul>
<li>Qualität von Content</li>
<li>Sichtbarkeit und Auffindbarkeit von Content</li>
</ul>
<h1>Generative-Engine-Optimization (GEO) im Sinne der Nutzenden</h1>
<p>Dabei wird die Sichtbarkeit und Auffindbarkeit von Content immer mehr zu einer Frage von AI Visibility: ob die eigenen Inhalte und die eigene Marke in den Antworten und Ergebnissen von KI-Chatbots und -Agenten stattfinden.</p>
<p>Ich kann dich beruhigen: LLMs priorisieren Content nach Kriterien, die auch für Nutzende relevant sind.</p>
<p>Neben der technischen Dimension – Content muss für KI-Systeme lesbar und verwertbar sein – muss Content dafür folgenden Dimensionen gerecht werden: </p>
<ul>
<li>Nutzerzentriert aufgebaut</li>
<li>Vertrauenswürdig</li>
<li>Klare thematische Positionierung</li>
</ul>
<p>Damit hängen die beiden Aspekte – Qualität und Auffindbarkeit von Content – stark zusammen.</p>
<p>Um als Team an diesen Punkt zu kommen, müssen wir zuerst die Grundlagen klären: Ohne eine saubere Ausrichtung ist kein Content-Team der Welt fähig, mittelfristig Content in dieser Qualität zu produzieren.</p>
<h1>Die Triple-Content-Line</h1>
<p>Unsere Ausrichtung als Content-Team definieren wir in folgenden Elementen:</p>
<ul>
<li>Content Strategie</li>
<li>Content Guidelines</li>
<li>Content Governance</li>
</ul>
<p>Als Team einigen wir uns dort auf den gemeinsamen Nenner unserer Content-Arbeit. Nur so können wir qualitativen Content, der internen und externen Ansprüchen gerecht wird, produzieren.</p>
<p>Dabei ist die Abgrenzung zwischen Content Strategie und Content Guidelines nicht ganz einfach. </p>
<h1>Was gehört in eine Content Strategie?</h1>
<p>Es gibt keine abgeschlossene Antwort auf die Frage, was in eine Content Strategie gehört. (Unabhängig davon würde ich eine Content Strategie auch eher als einen Arbeitsbereich beschreiben als ein abgeschlossenes Dokument.)</p>
<p>Für mich gehören folgende Elemente in eine Content Strategie:</p>
<ul>
<li>Content Ziele: Was will ich eigentlich mit meinem Content erreichen?</li>
<li>Zielgruppen und ihre (Content-)Bedürfnisse</li>
<li>Content Journeys der Zielgruppen</li>
<li>Kanalstrategie: Mit welchen Kanälen arbeiten wir und wie spielen diese zusammen?</li>
<li>Kernbotschaften</li>
<li>Erfolgsmessung</li>
<li>Schnittstellen zu Content Governance und Content Lifecycle Management</li>
</ul>
<h1>Was gehört in Content Guidelines?</h1>
<p>Content Guidelines müssen im Unterschied zur Content Strategie im Tagesgeschäft verankert sein. Neben fixen Orientierungspunkten (z.B. angestrebtes Sprachniveau) enthalten sie bei uns oft Prinzipien:</p>
<ul>
<li>Sprachniveau</li>
<li>Sprachregeln (z.B. einheitliche Schreibweise von Daten, Währungen und Ähnlichem, aber auch Themen wie Gendern oder Zielgruppenansprache)</li>
<li>Tonalität</li>
<li>Seitentypen</li>
<li>Tools</li>
<li>...</li>
</ul>
<h1>Was gehört in die Content Governance?</h1>
<p>Content Governance regelt, wer im Content-Team wofür verantwortlich ist – und wie wir zusammenarbeiten, damit Strategie und Guidelines auch tatsächlich gelebt werden. Ohne Governance bleiben Strategie und Guidelines gut gemeinte Dokumente, die im Tagesgeschäft untergehen.</p>
<p>Für mich gehören folgende Elemente in die Content Governance:</p>
<ul>
<li>Rollen und Verantwortlichkeiten: Wer erstellt, wer reviewt, wer gibt frei?</li>
<li>Freigabeprozesse: Wie läuft Content von der Idee bis zur Publikation?</li>
<li>Pflege- und Aktualisierungsprozesse: Wer ist wann für ein Update zuständig?</li>
<li>Qualitätssicherung: Wie stellen wir sicher, dass Guidelines eingehalten werden?</li>
<li>Eskalationswege: Was passiert bei Uneinigkeit oder Sonderfällen?</li>
<li>Schnittstellen zu anderen Teams (z.B. Design, Produkt, Marketing)</li>
</ul>
<h1>Das Ziel? Zusammenspiel</h1>
<p>Content Strategie, Content Guidelines und Content Governance greifen ineinander: Die Strategie gibt die Richtung vor, die Guidelines übersetzen sie in den Alltag, und die Governance stellt sicher, dass beides nicht nur auf dem Papier existiert. Erst im Zusammenspiel dieser drei Elemente entsteht Content, der konsistent, vertrauenswürdig und für Nutzende wie für KI-Systeme auffindbar ist. Und den Positionierungszielen des eigenen Unternehmens gerecht wird.</p>
<p>Content-Produktion wird günstiger. Aber ohne diese Grundlagen bleibt sie beliebig – egal wie leistungsfähig die Tools werden. Wer jetzt in Strategie, Guidelines und Governance investiert, legt das Fundament für Content, der auch morgen noch gefunden, verstanden und dem vertraut wird.</p>
<p><em>Wie Generative-Engine-Optimization (GEO) operativ funktioniert, folgt in den nächsten Posts.</em></p>]]></description>
    </item>
        <item>
      <title>Apertus after 8 months - what we learned and look forward to using Switzerland&#039;s AI Model</title>
      <link>https://www.liip.ch/de/blog/apertus-after-8-months-what-we-learned-and-look-forward-using-switzerlands-ai-model</link>
      <guid>https://www.liip.ch/de/blog/apertus-after-8-months-what-we-learned-and-look-forward-using-switzerlands-ai-model</guid>
      <pubDate>Thu, 09 Jul 2026 00:00:00 +0200</pubDate>
      <description><![CDATA[<p>Since September 2025, we have been running <a href="https://www.apertus-ai.org/">Apertus</a> — the Swiss open-source language model developed by <a href="https://www.epfl.ch/">EPFL</a>, <a href="https://ethz.ch/">ETH Zurich</a> and the <a href="https://www.cscs.ch/">Swiss National Supercomputing Centre (CSCS)</a> — on <a href="https://zuericitygpt.ch">ZüriCityGPT</a>, our reference implementation of a City AI assistant. Eight months later, I want to share what we have learned and what we are curious about for the upcoming Apertus 1.5 release.</p>
<!-- IMAGE: Screenshot of zuericitygpt.ch and oss.zuericitygpt.ch side by side -->
<h1>A brief history of LiipGPT and ZüriCityGPT</h1>
<p>2023, Chregu released the <a href="https://www.liip.ch/en/blog/ask-zuricitygpt-anything-about-the-government-of-the-city-of-zurich">first version of LiipGPT</a>. 2024, we were able to <a href="https://www.liip.ch/en/blog/zuricitygpt-oss-version-using-only-open-source-models">use only open-source models on ZüriCityGPT</a>. 2025, we switched the open-source version to the <a href="https://www.liip.ch/en/blog/apertus-4-ways-to-try-out-switzerland-s-new-ai-model">newly released Apertus model</a>. In this blog post, we talk about the insights gained since then.</p>
<h1>Setting up the experiment</h1>
<p><a href="https://liipgpt.ch">LiipGPT</a>, our <a href="https://www.liip.ch/en/work/projects/liipgpt">generative AI framework</a>, has a pluggable model layer. That means we can swap the underlying language model without changing anything else — same Retrieval-Augmented Generation (RAG) pipeline, same knowledge base of official content from stadt-zuerich.ch. This gave us a straightforward way to compare models in a real deployment.</p>
<p>The production system at <a href="https://zuericitygpt.ch">zuericitygpt.ch</a> runs GPT-4o-mini on Azure Europe. Alongside it, we run an open-source variant at <a href="https://oss.zuericitygpt.ch">oss.zuericitygpt.ch</a> where in the past we compared Llama and Mixtral models. We switched the open-source variant to <a href="https://www.swisscom.ch/en/about/news/2025/09/02-apertus.html">Apertus-70B</a> when it was released in September 2025, then moved to <a href="https://www.infomaniak.com/en/hosting/ai-services/open-source-models">Infomaniak-hosted</a> Apertus in April 2026. (To be clear: ZüriCityGPT itself and the LiipGPT platform are not open source — but the underlying language model on the OSS variant is.)</p>
<p>We run Apertus as the 70-billion-parameter model. It was trained from scratch on the Alps supercomputer using Swiss hydroelectricity, with 15 trillion tokens of training data spanning over 1,800 languages. It is fully open-source under an Apache 2.0 licence.</p>
<p>Over the last eight months, we collected over 39,000 conversations on the production side and 3,400 on the open-source side. Quality is measured using three metrics: how often an answer is acceptable (Match Rate), how often it is actually good (Match Rate +1), and how faithfully it reflects the source documents (Faithfulness). These are metrics our LiipGPT team has established as practical product metrics, not scientific benchmarks — the data is calculated real-time and asynchronously by the LiipGPT platform while users ask different questions over time, evolving website content, and model endpoint changes. Parts of the analysis were prepared with the help of Claude. Our aim is not to provide a scientific study but more a practical report. I share these results with the Apertus community because real-world deployment data of this open-source model is rare — and we think it's useful.</p>
<h1>What we observed</h1>
<p>Both models get the job done. On the basic question — does the user get an acceptable answer? — both models perform similarly. GPT-4o-mini scores 81%, Apertus 82% on historically matched questions. In a controlled test with 20 identical questions asked at the same moment, both hit 100%. Both models generally work as expected.</p>
<p>The gap shows up in answer quality. When we raise the bar to "is this answer actually good?", GPT-4o-mini scores 72% and Apertus 55%. That is a noticeable difference, and it matters for users who expect precise, well-structured responses.</p>
<p>Faithfulness is more nuanced than we expected. GPT-4o-mini averages 0.79, Apertus 0.50 on our faithfulness metric. But these numbers need context. We noticed something we started calling the "faithfulness paradox": when a model says "I cannot answer this," you might expect high faithfulness — no hallucination, after all. In practice, these evasive answers score low because the model is failing to use information that is available. A model that engages with its sources and cites them scores higher, even though it is making more verifiable claims. Both behaviours have tradeoffs, and both showed up across the two systems.</p>
<figure><a href="https://www.liip.ch/media/pages/blog/apertus-after-8-months-what-we-learned-and-look-forward-using-switzerlands-ai-model/ac2f708a50-1783608445/chart_comparison_2x3.png"><img alt="" src="https://liip.rokka.io/www_inarticle_5/b851b8/chart-comparison-2x3.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/b851b8/chart-comparison-2x3.jpg 2x"></a><figcaption>Three quality metrics compared across a controlled test (20 identical questions) and historically matched real-user questions</figcaption></figure>
<p>Sometimes Apertus gave the better answer. The most striking example was the waste dump question: "What are the rules for using the Zurich waste dump?" GPT-4o-mini responded cautiously, saying it could not provide specific rules. Apertus gave a detailed, structured answer with recycling centre details, phone numbers, and cited sources — scoring 6/6 on quality where GPT-4o-mini scored 3/6. When the knowledge base contains a clear, structured answer, Apertus can be surprisingly good at extracting and presenting it.</p>
<figure><a href="https://www.liip.ch/media/pages/blog/apertus-after-8-months-what-we-learned-and-look-forward-using-switzerlands-ai-model/9436173ed7-1783609912/apertus_gpt4omini_waste_dump.png"><img alt="" src="https://liip.rokka.io/www_inarticle_5/3b465f/apertus-gpt4omini-waste-dump.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/3b465f/apertus-gpt4omini-waste-dump.jpg 2x"></a><figcaption>GPT-4o-mini and Apertus respond to the same question</figcaption></figure>
<p>That said, this willingness to commit cuts both ways. In cases where the source material is ambiguous, Apertus can feel more helpful while being less strictly grounded — and for public-sector chatbots, traceability matters as much as helpfulness.</p>
<p>GPT-4o-mini is more concise and better at synthesis. Apertus tends to produce longer, more verbose answers. In a public-service chatbot, that is a real UX issue — users want a direct answer with clear next steps, not three paragraphs of context.</p>
<figure><a href="https://www.liip.ch/media/pages/blog/apertus-after-8-months-what-we-learned-and-look-forward-using-switzerlands-ai-model/b6516e51a6-1783610106/apertus_gpt4omini_verbose.png"><img alt="" src="https://liip.rokka.io/www_inarticle_5/677113/apertus-gpt4omini-verbose.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/677113/apertus-gpt4omini-verbose.jpg 2x"></a><figcaption>GPT-4o-mini responds concise, Apertus more verbose</figcaption></figure>
<p>GPT-4o-mini also handles factual freshness better: it returned Zurich's current population (452,421), while Apertus fell back to a 2013 figure. When information has changed or the answer requires connecting multiple sources, GPT-4o-mini felt more reliable.</p>
<figure><a href="https://www.liip.ch/media/pages/blog/apertus-after-8-months-what-we-learned-and-look-forward-using-switzerlands-ai-model/5ed0fada3b-1783610421/apertus_gpt4omini_population.png"><img alt="" src="https://liip.rokka.io/www_inarticle_5/997ce3/apertus-gpt4omini-population.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/997ce3/apertus-gpt4omini-population.jpg 2x"></a><figcaption>GPT-4o-mini answers with more recent population data compared to Apertus</figcaption></figure>
<p>Language switching is a weak spot. When we asked "Est-ce que tu parles français?", GPT-4o-mini responded fluently in French. Apertus declined, saying it could not answer questions about its own linguistic capabilities. In more recent tests we saw that Apertus followed the language instructions in better ways, so we are looking forward to see this further improve.</p>
<figure><a href="https://www.liip.ch/media/pages/blog/apertus-after-8-months-what-we-learned-and-look-forward-using-switzerlands-ai-model/0bbc6670a2-1783610288/apertus_gpt4omini_language.png"><img alt="" src="https://liip.rokka.io/www_inarticle_5/0387e4/apertus-gpt4omini-language.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/0387e4/apertus-gpt4omini-language.jpg 2x"></a><figcaption>GPT-4o-mini follows the language instruction, Apertus did not</figcaption></figure>
<p>The Apertus-70B IFEval score of 44% (vs 58% for comparable models) confirms that instruction following is an area where the model has room to grow.</p>
<figure><a href="https://www.liip.ch/media/pages/blog/apertus-after-8-months-what-we-learned-and-look-forward-using-switzerlands-ai-model/38a7a3a17b-1783609312/chart_monthly_cleaned.png"><img alt="" src="https://liip.rokka.io/www_inarticle_5/c6a834/chart-monthly-cleaned.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/c6a834/chart-monthly-cleaned.jpg 2x"></a><figcaption>Monthly quality trend from October 2025 to June 2026</figcaption></figure>
<p>The trend is positive. In October 2025, Apertus achieved a Match Rate +1 of 59%. By June 2026, it had climbed to 63% — the best month on record. That is not a dramatic leap, but the direction is consistent and it came without a major model update. Meanwhile, GPT-4o-mini has been stable at 63–83% throughout. The gap is narrowing.</p>
<h1>What we are looking forward to</h1>
<p>The Apertus team is preparing version 1.5, with improvements expected in multimodal support, agentic capabilities, tool calling, enhanced reasoning, and better training for regional Swiss languages. We plan to update ZüriCityGPT's open-source instance as soon as it becomes available and continue our quality monitoring.</p>
<p>What encourages us most is that the collaboration is flowing in both directions. After our <a href="https://www.ailights.ch/">aiLights talk</a>, the Apertus team shared our findings with their staff. This close feedback loop between research and practice is exactly what makes the Swiss AI ecosystem worth investing in.</p>
<figure><img alt="" src="https://liip.rokka.io/www_inarticle_5/3719ab/04.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/3719ab/04.jpg 2x"><figcaption>Christian Stocker presenting ZüriCityGPT. Photo: Oleg Lavrovsky (CC BY 4.0), source: <a href="https://log.alets.ch/115/">https://log.alets.ch/115/</a></figcaption></figure>
<p>My practical advice, for anyone considering sovereign AI: start with the use case, then choose the model. If your priority is maximum answer quality today, more popular models are usually the stronger choice. If your priority is data sovereignty, transparency, and independence — and your knowledge base is well-curated — Apertus already delivers. The question is not "sovereign AI or quality" but "how much quality gap are you willing to accept for full sovereignty?" And that gap is closing.</p>
<p>Check the <a href="https://docs.google.com/presentation/d/1Q792ccRiiiRk6zIlmCu7GkkjjZ6kiZJfbRk_CSBhxJ8/edit?usp=sharing">Slides</a> and <a href="https://www.youtube.com/watch?v=Fk0g--aWi1M">Recording</a> for further details.</p>
<h1>Acknowledgements</h1>
<p>I would like to thank the ETH Zurich, EPFL, and the wider <a href="https://www.swisscom.ch/en/about/news/2025/09/02-apertus.html">Swiss AI Initiative</a> community for making Apertus available as a public good. As Martin Jaggi from EPFL put it: "We aim to provide a blueprint for how a trustworthy, sovereign, and inclusive AI model can be developed." (<a href="https://ethz.ch/en/news-and-events/eth-news/news/2025/09/press-release-apertus-a-fully-open-transparent-multilingual-language-model.html">source</a>) Eight months in, we can see that blueprint taking shape and are looking forward to test the 1.5 model in practice as well.</p>
<ul>
<li>
<p>Thank you <a href="https://www.liip.ch/de/team/chregu">Christian Stocker</a> and the <a href="https://liipgpt.ch">LiipGPT</a> team for your dedication to build a practical, experiment-driven and scalable solution for AI chat and search.</p>
</li>
<li>
<p>Thank you <a href="https://www.linkedin.com/in/sabine-wildemann/">Sabine Wildemann</a> and the <a href="https://www.ailights.ch/">aiLights</a> team for hosting our talk.</p>
</li>
<li>
<p>Thank you <a href="https://log.alets.ch/">Oleg Lavrovsky</a> and the Apertus community for your feedback and support.</p>
</li>
<li>
<p>Thank you <a href="https://publicai.ch/">Public AI</a> and our partner <a href="https://www.infomaniak.com/en/hosting/ai-services/open-source-models">Infomaniak</a> for providing the infrastructure.</p>
</li>
</ul>
<h1>Build with Apertus: join Hack Apertus</h1>
<p>Want to get hands-on with sovereign Swiss AI? <a href="https://hackapertus.ch/">Hack Apertus</a> is a two-stage open-source hackathon series where teams build on Apertus to develop blueprints for sovereign infrastructure across the public and private sector. Liip is a partner — we are excited to see what the community builds.</p>
<p>Sign up at <a href="https://hackapertus.ch/">hackapertus.ch</a>.</p>
<figure><a href="https://hackapertus.ch"><img alt="" src="https://liip.rokka.io/www_inarticle_5/afc775/hackapertus-partner-liip.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/afc775/hackapertus-partner-liip.jpg 2x"></a><figcaption>Hackapertus Liip Promo Image</figcaption></figure>]]></description>
    </item>
        <item>
      <title>Ein souver&#228;ner KI-Chatbot f&#252;r den Staat Freiburg</title>
      <link>https://www.liip.ch/de/blog/ein-souveraener-ki-chatbot-fuer-den-staat-freiburg</link>
      <guid>https://www.liip.ch/de/blog/ein-souveraener-ki-chatbot-fuer-den-staat-freiburg</guid>
      <pubDate>Mon, 15 Jun 2026 00:00:00 +0200</pubDate>
      <description><![CDATA[<h2>Den Zugang zu öffentlichen Informationen erleichtern</h2>
<p>Das Portal fr.ch vereint eine grosse Menge an Verwaltungsinformationen: Verfahren, Gesetzgebung, Formulare, kantonale Dienstleistungen sowie Informationen der Gemeinden. Auf einer so umfangreichen Website die richtige Antwort zu finden, ist nicht immer einfach.</p>
<p>Du stellst deine Frage ganz normal, so, wie du sie auch einer Person stellen würdest. Obwohl <a href="https://www.fr.ch/de">fr.ch</a> in Französisch und Deutsch verfügbar ist, wie es sich für einen zweisprachigen Kanton gehört, können Fragen in jeder beliebigen Sprache gestellt werden. So erhalten auch Bürger*innen, die keine der beiden Amtssprachen sprechen, Zugang zu den benötigten Informationen.</p>
<p>Der KI-Chatbot identifiziert relevante Inhalte auf fr.ch sowie in der Freiburger Gesetzessammlung und erstellt daraus eine Antwort. Dabei verweist er auf die offiziellen Quellen, die für die Antwort herangezogen wurden.</p>
<p>Diese Transparenz ist zentral. Die KI ersetzt keine offiziellen Inhalte, sie erleichtert den Zugang dazu.</p>
<h2>KI im Zeichen digitaler Souveränität</h2>
<p>Die Datensouveränität stand im Mittelpunkt des Projekts. Für die Staatskanzlei war es entscheidend, die vollständige Kontrolle über die Infrastruktur und die Datenverarbeitung zu behalten.</p>
<figure><img alt="" src="https://liip.rokka.io/www_inarticle_5/b24742/digital-sov-de.jpg" srcset="https://liip.rokka.io/www_inarticle_5/o-dpr-2/b24742/digital-sov-de.jpg 2x"></figure>
<p>Der Chatbot basiert auf <a href="https://www.liip.ch/en/work/projects/liipgpt">LiipGPT</a>, unserer Plattform für generative KI. Unsere Retrieval-Augmented-Generation-Technologie (RAG) wählt für jede Anfrage die relevanten offiziellen Inhalte aus. Die Lösung wird bei Exoscale in der Schweiz gehostet.</p>
<p>Als Large Language Model (LLM) setzen wir auf Mistral AI, einen der führenden europäischen Anbieter. Dieses Modell wird in der Schweiz von <a href="https://www.infomaniak.com/de/hosting/unsere-angebote-cloud-computing">Infomaniak</a> betrieben, einem Vorreiter für nachhaltiges und umweltbewusstes Webhosting. Die Server des LLM werden ausschliesslich mit erneuerbarer Energie betrieben und setzen auf natürliche Kühlung, Wärmerückgewinnung, optimierte und wiederverwendete Hardware sowie lokale CO₂-Kompensation.</p>
<p>Diese Architektur reduziert die Abhängigkeit von grossen US-amerikanischen Plattformen und setzt auf einen schweizerischen und europäischen Ansatz für generative KI. Anfragen werden anonymisiert, und es werden keine personenbezogenen Daten an Dritte weitergegeben. Die gesamte Lösung bleibt unter schweizerischer Rechtshoheit und erfüllt die geltenden Datenschutzbestimmungen.</p>
<h2>Eine Zusammenarbeit mit starken Wurzeln in der Westschweiz</h2>
<p>Die Staatskanzlei Freiburg hat das Projekt mit einer klaren Vorstellung der administrativen und politischen Anforderungen geleitet. Unser Team, unter anderem mit Mitarbeitenden in Freiburg, entwickelte die Lösung in enger Zusammenarbeit mit der Staatskanzlei. So konnten wir optimal auf die tatsächlichen Bedürfnisse eingehen.</p>
<p>Diese Nähe ermöglichte kurze Abstimmungswege und einen starken Fokus auf die Nutzer*innen-Erfahrung. Neben der technologischen Umsetzung ging es darum, ein Werkzeug zu schaffen, das für die Freiburger Bevölkerung nützlich, verlässlich und verständlich ist.</p>
<h2>Ein erster Schritt hin zu einer zugänglicheren Verwaltung</h2>
<p>Seit dem ersten Rollout im März haben wir verschiedene Elemente weiterentwickelt, um die Qualität der Antworten kontinuierlich zu verbessern. Dabei profitiert unsere Kundin insbesondere von unseren automatisierten Evaluationslösungen, mit denen die Qualität der generierten Antworten fortlaufend überprüft werden kann.</p>
<p>Diese erste Version deckt die öffentlichen Informationen auf fr.ch sowie die Freiburger Gesetzessammlung ab. Der Inhalt beschränkt sich jedoch nicht allein auf das kantonale Portal: Er wird schrittweise um Informationen erweitert, die von anderen kantonalen Institutionen und bestimmten Gemeinden veröffentlicht werden. Mit jeder neuen Quelle nimmt auch die Bedeutung einer konsequenten Qualitätssicherung zu. Der Kanton Freiburg unterhält übrigens <a href="https://www.fr.ch/de/der-ki-chatbot-von-frch">eine eigene Seite zum Chatbot</a>, auf der unter anderem die verwendeten Datenquellen aufgeführt sind.</p>
<p>Eine Überzeugung begleitet uns dabei: KI im öffentlichen Sektor kann souverän, transparent und nützlich sein. Voraussetzung dafür ist, die Kontrolle über die Infrastruktur zu behalten und die verwendeten Quellen jederzeit offenzulegen.</p>
<p>Möchtest du dich über KI in deiner Institution <a href="https://www.liip.ch/de/team/thomas-denervaud">austauschen</a>?</p>]]></description>
    </item>
      </channel>
</rss>