返回知识库
0

title: "在边缘运行 AI:直接在浏览器中运行实际工作负载"
source_url: "https://www.infoq.com/presentations/local-ai-browser-inference-privacy/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global"
author: "InfoQ"
excerpt: "公共早报 该演示探讨了在浏览器中本地运行 AI 模型的优势和挑战,重点介绍了隐私优先的转录、实时视频处理和交互式数据分析等用例,并讨论了本地推理与云端推理之间的权衡。"


演讲记录

James Hall: 我将给你们一点关于我自己的介绍和背景,只是给你们一些 context 关于我所做的各种事情。我要谈谈,为什么要费心?本地运行东西有什么意义?我将 all this with 一堆演示 so you can actually see and feel 这一切目前的样子。我将尝试 high level 解释这一切 stuff 的工作原理,并希望给你们 pointers 当你们应该使用它时。我还要给你们 plenty of pitfalls and takeaways。我们有很多 battle scars。我们犯了很多错误,所以希望你们能从我的错误中学习。

背景

稍微给点背景。我 15 年前创办了一家技术咨询公司。我们基本上为客户构建数字产品。我们收到了各种各样奇妙非凡的 brief,从街道照明到远程解锁汽车再到 AI 仪表板。这里面有很多丰富的东西。我的背景,我一直用 JavaScript 做各种奇妙的事情。只是给你们一点 sense 我有多喜欢 bending browsers out of shape。2012 年,我移植了一个 SPC700 模拟器。SNES 内部的声芯片,你可以 actually emulate that。这是非常早期的 Emscripten 日子,所以你可以移植 C++ 到 JavaScript。Lots of playing around with Uint8s 和 massive arrays of numbers。我实际上用当时非常新和闪亮的 Web Audio API 使这工作,是的,我 actually surprised myself。Most people would look at this kind of project and go, what's the point?

找到事物的边缘真的是好的,像这些 kinds of experimentation projects。2009 年,甚至更早,我写了一个叫 jsPDF 的库,它 haunted me to this day,并且最近实际上有 massive spike in downloads。你们只能假设是人们在 vibe coding invoice generation software,或者是 God knows what。是的,每天数百封电子邮件,很多开放的 PR 请求。I was automatically closing PRs way before people got grumpy with AI。这带来了很多有趣的项目和对话。我很高兴我写了它,但我讨厌 PDF。我也是 AI 回收项目的外部顾问。

另一个顾问是英国时尚纺织协会的负责人。我们正在试图弄清楚如何利用自动分拣减少纺织品浪费。我们长期以来非常热衷于 AI 及其应用。从街道照明项目开始,我们做故障自动检测和预测性维护。我们真的很喜欢当我们能真正弄清楚数学和实际应用如何映射到现实世界时。我们一直在举办活动让人们走出来,谈论不同的应用。

我第一次接触 LLMs 实际上是为一家初创公司工作时。我们为他们制作了一个网络浏览器,用 Kotlin 和一些 React。我们实际上在网络浏览器中嵌入了一个 LLM 来做自动标记和分类。这是在 ChatGPT 出现之前。我们不得不去 OpenAI 说,这是我们正在构建的应用程序。这是它做的。这些是我们看到的风险领域。他们非常谨慎地批准每个用例。显然,ChatGPT 推出后,闸门打开了,他们真的不再在乎了,可能是因为钱。这是一个非常有趣的过程。我们在这过程中学到了很多。我们还推出了一个基于 LLM 的合规平台。这赢得了年度技术创新奖。我们对聊天机器人了解了很多。我现在也知道了很多为什么你不想构建它们的原因。是的,很多 battle scars from this。

我们最终构建了很多现在司空见惯的系统。我们最终作为产品的一部分编写了我们自己的 LLM 评估和可观测性平台。我们还发现用户与聊天机器人交互,因为它们如此 open-ended,你需要将它们 into a railroaded experience。我们想到了 checklisting 的想法。我们有自己的 structured outputs 版本,这非常原始,试图让人们 down a certain path。我们在这过程中学到了很多。我们被英国一家大型博彩公司 picked up。我们正在与 Timeform 团队合作,将 ML 模型与 LLMs 混合。我们实际上接管了亚马逊网络服务做概念验证。在那里,他们无法 confidentially predict 马的顺序到任何程度。我想他们的 F-score 大约是 0.5、0.6,比随机好一点。

我们能够在几周内通过混合各种不同技术将其提升到大约 0.9。我们通过那个项目学到了很多。This is where a lot of this talk's content is coming from 是我们正在合作的一些新组织的一些想法。其中很多是工业物联网。很多数据。我们如何理解它?我们如何绘制它?我们如何 measure what the trends are?这给你们一点背景,关于这一切 framing 来自哪里,我的 brain is sort of at。

本地与云端

我想和你们谈谈本地和云之间的区别。我认为今天很多人会默认为云,因为这是最简单的方法。它有最强的模型,是最小阻力的路径。很多人选择这些提供商,sign their agreements,并将所有数据发送给这些公司,包括他们的客户数据。这种设置的问题是最终用户越来越多地开放更多和更多的个人信息,而没有真正考虑后果。我对组织内的日志记录平台有点担心。员工看数据的 checks and balances 是什么?如果被美国政府强迫 open up logs,例如,我们真的能信任公司做他们说的要做的事吗?

你有什么控制权?这就是让我开始有点担心的地方。这就是 server-side inference 的缺点所在。它不仅仅是隐私问题。还有低和无互联网连接状态。你们会很惊讶在很多国家,Wi-Fi 非常 patchy。如果你正在火车上,如果你 just out of reach of signal。还有网络延迟,这对于一些应用程序来说相当 jarring,特别是 around real-time video and audio processing。例如,如果你正在 cut out background noise,你不会想从英国到爱尔兰的数据中心然后 back out to your end user。成本随用户 scale,所以你会成为自己成功的受害者。假设你构建了一些超级受欢迎的东西,it relies on the latest and greatest OpenAI model。你得到一百万用户,你现在有一百万人要支付他们的推理费用。

我要向你们展示一个接近人类水平的 transcription quality demo,可以在 JavaScript 中使用。Whisper 模型最初由 OpenAI 团队生产,他们开源了权重。人们一直在 improve and riff on 这些模型在过去四年左右。它们越来越强大。它们能够有正确的标点。他们能知道地名。它们可以在特定行业做复杂术语。所有这些都可以在设备上、在网络浏览器中发生。我们一直在看到这方面的例子。我和 Ian 的好朋友之一,实际上 Nick Payne,一直在构建一个应用程序,我会演示,它使用本地转录模型 to make it a totally privacy-aware meeting notetaker。我构建的一个演示 to show off some of the capabilities of local models,这是几年前,在 Claude 能够 one-shot this 之前。

这实际上获取本地模型,它能够 look at the log probes。你可以对这些做数学以获得概率 so you can actually see the likelihood of each token that is happening。很多商业模型提供商 actually hide this information 为了阻止人们 reverse-engineering 他们的前沿模型。我认为一些 API 现在确实 expose 这些。使用本地模型的一个真正好的优势是你可以访问所有这些 raw data。你可以 actually see if all of the log probes are pretty flat,没有 standout,你知道 there is a risk of hallucination there,因为它 literally not able to pick。这是 talat。我和 Nick 坐在一起,reverse-engineered an app called Granola,这是一家伦敦的 startup。他们使用了一个非常有趣的 Mac API,允许你 tap into the audio interface。

他们也有一个很好的巧妙技巧,你可以弄清楚会议何时开始和结束。Nick 获取了一些这些学习并将其 into an open-source Taps library,你可以轻松 hook into the audio interface。这使用本地模型的真正良好 blend。这是一个 Electron app。所有 UI 都是 HTML,JavaScript in React。然后它实际上 hook into 这些本地模型,可以在 Apple Neural Engine 上本地运行。你有 NVIDIA's Parakeet。你有一个流式模型。还有一个用于摘要的阿里巴巴模型。然后我们也有语音指纹和识别模型。这些真的很好 can be quantized without losing lots of impact,已经达到 reasonable size to download。

当涉及本地模型时,不再只有一个选项。Until a couple of years ago,唯一的选择是 bring your own API,所以获取所有权重,compress them as much as you can。Quantize them。将它们从 8 位整数转换为 4 位或 2 位。Send them over the wire and cache them in the browser or in your Electron app or in your desktop application。那曾经是唯一的方式。那个整个生态系统已经走了很长的路。仍然有一些问题。Chrome 团队正在尝试解决和整理的问题之一是 same origin caching limitations。一旦我们获得了 ratified for this 的标准并标记,你们可以看到 websites or web applications using one model in one place,你们将能够 reuse it in other places。

那是自带 AI 路线。对于那条路线,你使用 WebLLM,构建在 WebGPU 上。还有来自 Hugging Face 的 Transformers.js。它们使用 ONNX Runtime。然后你有其他本地推理提供商,比如 TensorFlow。内置 AI 非常新。它最近之前相当糟糕。我 early private beta list for this in Chrome 当他们有一个非常 early version of Gemini Nano 时。我尝试构建一个浏览器扩展 just extract the recipe from a recipe page and get rid of the whole story and everything。你真的必须与之斗争 like the prompting,你必须真的 try to get it to do the right thing。它已经从那以后 leaps and bounds。那些 Nano Gemini 模型现在变得好多了。他们还开始 shipping 一堆 other inbuilt APIs 非常有意义。

不仅你有你的 prompt API,你现在有一个翻译器,一个 summarizer,和一个语言检测器。他们正在试图让 everyone on board。Chrome 和 Firefox 和其他组织成员都 align on the same kinds of APIs。这超级有趣。好像确实遗憾的是不 explore these routes 当你口袋里的硬件越来越先进时。你有更强大的神经处理单元。你有 specialized GPUs 现在在你的设备上运行用于 AI。好像确实遗憾 to just ship everything off to Sam Altman 当你可以 actually weigh up the pros and cons 并找到其他路径时。

我爱 Transformers.js。我认为这绝对是 magic。这里有很多非常酷的模型,有很多方式你可以与它们交互。它是 JavaScript 原生的。它将完全像 plain JavaScript 在 CPU 上运行,但然后它也能够优化自己。Back when I did like that js-snes-player,没有任何优化。你在每个周期内可以做的计算量非常有限。当 Wasm 和其他技术开始采用,那真的加快了你可以的速度。现在,有直接访问 GPU 通过 WebGPU,你可以 run at near-native speeds 现在用于推理,which is super impressive。显然,一切 just stays on the device。没有 API 调用,没有第三方处理。现在有 new backends 出现。WebNN 仍然非常早期,但它将能够在你的 Android 和 iOS 手机上非常专门的 GPU 上运行。

当你在看一个 machine learning model to run on-device,compare the sizes 非常有用。一些更小的模型会 be extremely good at writing like a SQL query 或者 something super basic like that,或者摘要。当你越来越大,它们对越来越复杂的任务更好。在我们看到的趋势中,人们正在想出越来越多聪明的方法来 pack these models down,像使用不同区域、不同层模型的 variable width quantizing。这超级有用。机器学习中有一个常见短语,但我想提醒你们,all models are wrong, but some are useful。我认为这是非常重要的,当你用 AI 开发应用程序时,we should not be trusting the output just on face value。我们应该测试并找出我们在什么可接受的 bounds we find this model useful,以及它对我们的用户安全使用吗?Setting the goalposts up and setting up your evaluation suite at the start,这就是大部分工作所在。实际上集成模型是 easy bit。难度来自于围绕它的一切,so testing, validation, security。

有很多东西你可以做 to optimize in-browser inference。我不会 recommend quantizing your own models,除非你真的知道你在做什么,但你可以改变 precision。对于很多应用程序,它实际上并没有 lobotomize it as much as you might expect。你可以 take a 7-gigabyte model down to 2 gigabytes,而你只会有 modest quality loss。如我提到的,有 WebGPU 用于硬件加速,还有 Wasm 用于通用 CPU fallback,which is super handy。然后 WebNN 东西在接下来一两年出现。你也可以 fuse kernels together,你可以有 custom operators。你可以有 pre-compiled routes。有很多 model-specific graph optimizations 你可以做,还有 caching。WebLLM 在暴露其中一些功能方面超级好。你实际上可以选择你想要使用的模型。

你可以 actually test out whether it's going to work for your use case。我希望的是,当人们开始 more and more 使用这些东西,and the cross-origin caching comes into play,你会发现你可能最终拥有一台配备 Llama 3-point whatever 的台式机,and you can rely on it being there。目前,我们 not quite there,但你需要 skate to where the puck's going。这是我超级兴奋的。这是给我们那个 raw access to the NPU acceleration 的项目。这让我们达到即使在移动设备上也能达到 super-fast speed。这是我们正在寻找的 tipping point 当我们要构建这些 privacy-conscious applications on top of web technologies。我的一位客户在高度监管的空间工作。他们不希望工业物联网信息进入商业前沿模型。

实际上有一个很好的事业和 lots of interesting experimentation 来 actually run all this locally。有一个非常好的 super-fast DB 叫 DuckDB。他们有一个 Wasm build,出奇地快。我们一些客户 surprised by how quickly taking large datasets from Redshift,where the analytical workload you'd expect to run fairly quickly,实际上在本地设备上用非常优化的 DuckDB 看到了更好的结果,which is super surprising for them。如果你用 Parquet 编码它,你可以让它在浏览器中本地运行。这真的是为 try lots of permutations of queries 来 find insight within data。你可以用 LLM as a mini data scientist,尝试很多不同的角度,seeing the outputs and iterating very quickly。我们一直在用 virtualized CLI 做这个,which I'll talk about。

另一个 extremely good use case for local AI 是使用非常小的模型,像这个 token classification model。有一个 named entity recognition model,which is super small,它可以在 plain text 上运行。它可以拉出来,not using regular expressions or anything,但它可以拉出来 like this looks like a person's name, this looks like a location, this looks like an address。一旦你把所有那些 pieces 拉出来,你可以 swap it out so you're not sending it back to the server,之类的事情。或者你甚至可以显示警告, saying it looks like you're inputting patient data。当我们做合规应用程序时,因为这是人们工作的非常敏感领域,even if you instruct people not to upload sensitive information or documentation,often you'll find users, the first thing they do is copy and paste a load of patient data or some clinical study material into your UI.

这就是像这样的东西会超级方便的地方 because you can just alert them straight away, saying this data can't be processed for these reasons。我们实际上有 really good browser support now,which we didn't do a couple of years ago。终于,Safari has come out the gate with extremely good support。Firefox 第二,Chrome 首先。它真的 now getting really good。基本上 every other browser is secretly Chrome,所以这是 nice and wide support。这些模型 not just process text。它们可以 do lots of different applications。这是一个超级有趣的例子。我认为这是 Meta 项目,实际上。它来自 Meta Labs 项目之一。你可以 actually generate music based on just text alone and emotions。你可以说,yes,generate me an electro pop tune,它会 churn out a nice little piece of music。它变得 surprisingly good。

正如我一直在说的,in-browser inference 有这些非常大的好处,但它并非没有陷阱。你想要 measure which workloads make the most sense。对于某些工作负载,shifting that inference cost from your service to the user's device makes lots of sense,特别是 video-heavy applications,你可能使用 WebRTC,它是点对点的。You don't want to be bouncing people through a server。有一些非常好的例子。你可能用过 Google Meet。他们实际上使用一个非常简单的模型来做客户端的 background removal,which is super nice。你可以 actually just figure out for which workloads are you going to do these hybrid approaches?哪些你要 completely local?哪些你要 server-only?目前有点 catch,这些 larger models 是更复杂的推理。更好的匹配仍然是服务器端。真的是 making sure that your workload matches the kinds of models you want to run。

你应该衡量什么?

正如我提到的,deploying a model 的大部分工作是 measurement and making sure it meets your criteria。你需要衡量什么来弄清楚你想在浏览器内 launch 还是云端?你想要 look at time to first token。这用于 LLM。spin up the model 需要多长时间,warm it up,并生成第一个 token? Surprisingly for really small models it's super quick locally,实际上 that round trip to the server 是相当大的一块。特别是如果你发送 tens, or hundreds of thousands of tokens up,你想 think about overall latency and download times。如果这是一个 meeting note transcription app that someone's going to use every day,are they happy downloading 500 megabytes?可能。It's about doing those calculations and making sure the fit is right。

你想要 look at your token throughput。它实际上 output and do completions 有多快?然后,最难的是准确率。你也可以使用 LLM-as-a-judge。我总是 recommend that if you're able to do something using dumb code then do that。Not everything has to be AI。我经常看到很多 AI orchestration or evaluation suites lean towards using yet more AI to test the AI,并不总是 necessary。

通过架构而非策略实现隐私

通过在边缘设计带有 AI 的产品为用户做的伟大事情之一是你实际上 enforce privacy by your architectural decision rather than policy。正如我之前提到的,你不知道会发生什么 with let's say there's a big cloud provider like AWS。他们能被 compelled to start tracking information,开始 surfacing different log files 之类的事情吗?你可以 actually circumvent all of this by making it impossible for that to happen。Why don't you keep the things that you really need to be local, local。我要再次谈论 meeting note example。很多 meeting note taking apps 喜欢 Granola,像一些其他的,they basically send raw audio data from your device straight up to the cloud。其中一些提供商 actually have had breaches where API keys have been accidentally published。

Just the fact that they're logging and processing this even just for a moment,it's almost like the toxic waste of user data processing。就像你不想 handling this super sensitive information。Similar problems exist with the Meta Ray-Ban glasses。你一直在 recording all the time。它去某个数据中心,或者人类 potentially reviewing it to figure out problems,tagging,做所有那些 stuff。你有多开心 that's happening?你的用户有多开心 that's happening?当谈到 adding AI to your product 时,what I would wholeheartedly recommend 是稍微退后一步,please try not to just add a chatbot to your product。我这样说的原因不是因为我不认为自然语言处理 style features aren't cool。我确实认为它们是酷的。我认为有很多原因。

首先,for users who are not as patient as your internal developers or stakeholders,它们不太愿意花五分钟 chatting away trying to figure out what this bloody other chatbot is actually good at doing within your product。他们过去有过糟糕经历,所以 they might have gone to their courier,DHL,之类的。他们 used an online support chatbot and had mixed results。人们正在经历 chatbot fatigue。我不是说不要为你的产品添加一些自然语言处理功能。如果你想要绝对可以。如果你的 first go-to isn't let's bolt on the chatbot。

常见陷阱

你有责任,你有 technical power to figure out like, what is an LLM?What is a model?什么是我们正在构建的这个软件 actually good at?为你的用户做那个选择。Make it super easy。如果你正在构建一个 AI dashboarding product,the first thing you present them is not ask it anything。那太 open-ended。用户不够有想象力,但可能是 power users。Ninety percent of people just want to get in, get the job done 然后继续。Why don't you find out what it's good at?Why don't you present the suggestions?为什么不已经在后台处理了 these interesting insights and showing examples of what you could produce?然后如果你必须,你可以添加自然语言调整和建议,"Let me change this chart. Let me do this. Let me do the other."那是给 power users 的。

好的。留在那里。不要把它 like front and center。如果它不需要 AI,不要让它成为 AI。我知道这听起来愚蠢,但我们与很多客户合作,at the top of their agenda is we need AI,实际上他们不需要。这 what's wrong with a well-crafted SQL query or a regex。Just reach for a model when the problem is generally difficult and fuzzy,并不只是因为它可用。另一个本地优先的陷阱是 that first download is quite painful。你想要 aggressive caching。你想要显示进度。Ideally,fold it into the flow so they don't even notice。Instagram 刚出来时,每个人都超级惊讶 because they couldn't figure out how it was so fast。你会拍照然后 choosing your filter and you press go and it was basically there。

这是在我们真的有糟糕移动网络的时候。用户界面超级好。这实际上是一个简单的技巧。All they did was,as soon as you took the photo and went to the next step,it started uploading the photo,and in the background as you're choosing the filter and press send,all it sent was the text name of the filter to the backend and did the processing remotely。这是一个非常好的方式,if you can hide loading time in some other process,perceived loading time is all that matters。用户只关心他们对重量的感觉,而不是实际重量。

数据科学中的一个常见笑话是,你不需要 AI,你只需要一个好的 SQL query。我想稍微翻转一下,AI 的一个非常好的用例实际上是 write SQL queries。正如我在演示中展示的 DuckDB 例子,你可以 actually run this full analytical engine in a browser。你可以找到一些有趣的角度。我们一直在做的是通过在浏览器中运行一个假的 CLI 并允许 LLM write command line calls and arguments and pass SQL and do things like jq selectors 来 find interesting angles and insights in data using the fewest tokens possible。还有更小、更快、专门用于编写 SQL 的模型。一个快速方法 to figure out if you should be reaching for in-browser inference 是,is this data sensitive or not?Could it potentially include patient data?

Could it potentially include PII?如果是,有非常好的理由使用本地推理。即使数据最终进入云端,你至少可以 redact the PII 如果你不期望 store PII。再说一遍,就像有毒废物。你不想让它进入你的日志文件。你不想让它进入你的 S3 buckets。除非你的工作是处理那个 PII,you'd rather not have it anywhere near your servers。就不要使其成为一个选项。如果数据不敏感但你有高质量或复杂性要求,那可能是使用云 API 的好用例。如果你能找到足够好的模型在本地做,那就本地做。如 always,如果它不需要 AI,不要使用 AI。

模型在行动

我要展示一些我认为真的很酷的模型。我们这里有一个 audio multimodal demo,我拍了一张图片。示例图片是一些面包,但这是一个相当酷的小演示。它获取 2D pixel map,它能够推断那些像素的深度。我还 here taken a test image,这是 QCon logo 的一个小图片。这是一件相当酷的事情,you're actually doing some transformation into WebGL。这 actually coming on really quickly。你实际上能够做非常有趣的事情 like frame-by-frame video tracking。这非常 processor intensive,所以我们稍微加快了。它正在制作所有在屏幕上移动的东西的地图,and then you can actually build an interactive UI to then just select the flags,and it'll actually follow it around as the camera pans。

超级有用的 for blurring things out, highlighting things。是的,非常有趣的东西。我认为本地模型的一个非常好的用例是模糊背景之类的事情。人们经常在家工作。他们有孩子在后台跑来跑去。他们有很多 mess 他们没有收拾。你可以 actually do this super quickly now,like do real-time background removal。为什么不在你通过 WebRTC 发送之前本地做呢?因为 WebRTC 是点对点视频会议解决方案,没有理由你应该 up to the cloud and back out again,特别是如果这只是两个人之间的个人电话。

我要展示一些基于文本的模型例子。这是 Chrome 内置 prompt API 的演示,它实际上是 pretty decent。它并非总是如此,正如我提到的。基本上 one line of code,你可以设置一个 prompt。你可以传入文本,它可以完全在浏览器内完成 completions。没有任何网络调用 whatsoever。The model's been trained on QCon's website,显然,因为它知道你应该来 QCon 的所有好理由。它 also able to do constrained sampling and structured outputs,所以你可以 actually get JSON and stuff from it。是的,超级方便。列出所有参加 QCon 的最佳事物 as a short JSON block,然后你就有。你没有给 Sam Altman 任何钱,你有一堆 JSON 然后你可以 render into a nice UI。我总是 recommend,if you can,try not to show users walls of text。

如果你使用 LLM,总是 try and structure it in some way and make some nice UI elements。它只是让人们更容易理解,而且 they've got that wall of text fatigue。这是一个叫 Duck-UI 的产品。你可以免费试用。这完全在浏览器内。我在这里加载了几年 QCon 演讲的 CSV 文件,你可以 just ask it questions about the data,and it'll generate the SQL in WebLLM,所以它在你的 GPU 上运行。然后查询在 Wasm 中运行并返回数据。这些数据都没有离开你的网络浏览器。这些数据库适合 between 100 and 500 megabytes 的数据集大小,之类的东西,你可以 fit a surprisingly high amount of data in。

衡量什么是好的

推出模型的重要事项之一是 trying to measure what good looks like。你可以提出和 whiteboarding 尽可能多的指标,但很多利益相关者会专注于 does it actually improve a metric that they care about?可能是 like time to decision,或者是 decision quality,或者是 how closely it matches what a human would do。每当我们做项目时,我们会设定那些 criteria at the start,and we'd like to save as much of it in a database as we can,so that we can create this eval suite around the model。你如何衡量 LLM 输出的质量?在 regular software that isn't using LLM,it's deterministic,你只是 wrap automated tests around things。对于 LLMs,更狡猾的是 output。它更长,可以变化,and it gets quite tricky。

Instead of manually reviewing all the reasoning output from an LLM,那应该是一个机器人版的 Judge Rinder,你可以 actually get another model that is better to actually rank the thinking,并 rank whether you think that was a good output or not。你希望 judge model 比被你评判的模型更好。你会使用前沿模型来评估和测试较弱 in-browser model。我 wholeheartedly recommend 对于 any AI project,cloud or desktop,构建某种 eval suite 非常 visual and very easy for subject matter experts to understand。如果你正在做 like banking KYC,或者你在做 healthcare records,无论它是什么,你需要构建一个 UI,那个 healthcare person can understand,他们可以 see what the AI is doing。你不希望这是 an engineering thing。

这是 80%,90% 的你的产品。如果你正在构建一个非常 AI heavy 的产品,大部分产品将在 testing 中,以及如何 communicate that testing to people who know much better than you do about whether that's a good quality output or not。例如,let's say you're building a tool that can figure out text on a passport。也许它 quite low risk。也许这是 for checking in to a flight or boarding。你会存储 test fixture。这是 test file,the scan。Ideally,这是 like the awkward passport scan that's not working so well。然后你会有一个 JSON blob,which is,这是我们所期望的,然后你会比较 that to what actually happened。你想要 produce 的用户界面是这样的。你有,here is your fixture,here is what we expected,and here is what actually happened。你想要 make it super visible like this 的原因是因为你要让人们 pull various levers to improve the quality here。

你能 pull 什么 levers?超级简单的之一,改变底层 prompts。Obvious for an LLM。你会惊讶有多少产品团队进入 who have no way of knowing if changing a prompt makes their product better or worse。他们只会 ship it based on whether it felt a bit better,which is super dangerous because it's absolutely random。这是 super important that when somebody changes a prompt,you're able to rerun all of your edge case tests and figure out,have I made it better or worse?它听起来很明显,但很多人错过了这个。你可以 obviously change the underlying models used,所以 let's beef it up to a bigger model。让我们看看会发生什么。你也可以改变设置该模型的 harness,so the workflow。我不太喜欢术语 agentic,但 like an agentic flow where an AI 能够实际经历 multiple steps in order to produce a result。

我认为它需要尽可能 guardrailed。有一个惊人的项目叫 just-bash,这是 Vercel Labs 团队的。他们 actually engineered 一个完全 in-browser Bash environment,用 TypeScript 完全编写。你可以 write new commands in TypeScript,这将运行你想要的任何代码。然后 LLM 可以使用非常 short, succinct CLI-like syntax 将那些命令链接在一起,which is super good,因为使用 MCP,你在做很多大 JSON blobs,很难链接在一起,and you're wasting loads of tokens。使用这些虚拟化 Bash 环境,你可以让 LLM iterate and try things out in order to get to a result, much like Claude Code does。人们尝试各种花哨的东西 like indexing entire codebases in vector DBs and doing all this fancy shit,但当它 down to it,LLMs 实际上只是 really good at just using cat and grep and head and stuff。

为什么不 just let it use the really simple, dumb tools that it's really good at doing,but in a completely sandboxed way?LLM 不知道这不是真正的 Bash。它没有操作系统权限。它无法访问网络,例如。LLM doesn't care。不要做某事 super dangerous and connecting the LLM to your real CLI,why not connect it to a fake CLI?一旦你 pull those levers,你可以 rerun your lovely, nice and visible evals,and your subject matter experts can tell you whether they think it's suitable to ship or not。在我看来,我认为 all of the roads lead to madness,这是仅基于 vibes shipping alone。

策略和关键洞察

我要以一些策略结束。虽然我一直谈论的很多是 about picking models 以及更多技术方面,很多 actually improves AI apps 是超级无聊的。就像 talking to end users, talking to subject matter experts。就是 making sure that the platform is reliable,so it retries things when it fails。就是 preparing and describing data properly。就是 optimizing that end-to-end workflow with feedback,so that eval suite,so you can visually see and test things。是的,就是 lots of little tweaks,reviewing and evaluating over and over。我想要 today 带走的,如果数据是敏感的,去本地。如果是某人的音频或视频在家里,then try and apply some pre-processing to make sure you're not uploading content that you don't want to be uploading。你想在它上传之前 remove PII 吗?就做吧。Make sure you benchmark it on your actual workloads with the actual edge cases that you're going to be seeing in the wild。Have fun and build some cool stuff。Transformers.js 是一个非常好的起点。Just Bash 超级 fun。我们建议 give it a whirl。

查看更多带转录的演示文稿

AI知识库 / 在边缘运行 AI:直接在浏览器中运行实际工作负载 0 字 0 行 iliuqi
2026-09-04T08:58:42.281453193Z 2026-09-04T09:28:33.429665046Z