|
Jev has been all over the Internet for the past week. People are obsessed and declaring it an entirely new way to do AI. Is that hyperbole? Yes, definitely. But is it still very interesting? Absolutely. What is Jev? In more technical terms, it's a zero-shot generalizable classifier. In plain English, Jev is different from a large language model because it doesn't generate words. It makes decisions from the possible answers you give it, with a probability for each answer. We've had machine learning classifiers for many years to label data and help us categorize information. Jev is different because it:
We tried Jev on 499 of my email threads. The question was simple: which messages actually need my attention? The first version gave Jev a short snippet and asked it to pick between ten different labels I use to sort my email by urgency and type. That turned out to be a complete failure. It made 7 decisions out of 499. It failed because no one label had a high enough probability to be chosen! I needed to rethink the plan and provide more context to support the decision. In the second pass, we gave it more of each message (with a cap) and broke the job into smaller yes-or-no questions. The revised run analyzed and labeled all 499 threads correctly in about two and a half minutes. I confirmed this by comparing its output with my previous LLM-based labeling. Here's the 30-second look at the flow: Silent 30-second animation: Jev classifies an illustrative email, an LLM drafts a reply, and I review it. Jev surfaces the messages. A language model drafts replies. I review and edit them, then decide what to send. That zero-shot generalizable classifier handles one decision in my workflow. The rest of the system matters too: who gets to change an agent, what a finished job costs, and how we check the agent's work. 1. Who gets to change the company agent?What happened: Cloudflare's MCP server portals reached general availability. A company can put approved MCP servers behind one endpoint, manage access centrally, and log activity. Okta made Agent SSO, Agent-to-Agent Connections, and Resource Access Certifications generally available. Agent SSO gives Okta customers short-lived, identity-governed tokens for agent access. Agent-to-Agent Connections sets rules for which agents may call each other and records the handoff. What changes: Imagine Marissa in marketing with a social copywriter agent whose drafts are a little stiff. She wants to update its instructions and examples. That's one decision: who can change the agent? Cloudflare OS lets people build and revise personal apps, then share blueprints others can edit as their own copies. In Block's Buzz, an agent owner can edit its instructions, model, and who may direct it. The second decision is what can the agent access? Cloudflare's portal gives Marissa an approved list of MCP servers without making her configure every connection herself. Okta can govern which systems her copywriter reaches and which other agents it may call. For example, the copywriter could ask a company knowledge agent for current information from the wiki, if that connection is approved. Okta already lets admins deactivate an agent to block new sessions; it is working on an expanded kill switch to revoke active tokens and shut down sessions in progress. Marissa should be able to improve the copywriter's writing without waiting for IT. IT can decide separately whether it gets access to the publishing account, customer data, or the company knowledge agent. That's how I approach agent teams at DianoAI: the person who knows the work should be able to make changes to the agent. Access and changes should be visible. What to do next: Try Marissa's test on your agent platform. Can she change the copywriter's instructions herself? Can she find the approved tools and request access to the company knowledge agent? Can an admin see and revoke those connections without having to edit her copywriter? 2. A cheaper token can still mean a more expensive jobWhat happened: On Tuesday, OpenAI released GPT-6 Sol and Luna, with standard short-context API prices starting at $2/$10 and $0.10/$0.50 per million input/output tokens, respectively. Anthropic released Opus 5.5 at $4/$20. What changes: Those prices don't tell you what it costs to finish a job. The model may need more output tokens, retries, or a person's time to correct it. In one independent Grok 4.7 versus 4.6 evaluation, the token rates were the same, but the newer model cost more per evaluated task. What to do next: Run the same small set of real tasks through each model you're considering. Record the cost of an answer you can use, including retries and review. 3. Google open-sourced a security review harnessWhat happened: Google open-sourced Mantis, a security-focused harness for coding agents to find, reproduce, and patch vulnerabilities. It also described an internal pipeline that checks infrastructure code changes with agents and human review. Google says it prevents hundreds of vulnerabilities a month from reaching its code base or production. What changes: Development teams can try this kind of security review on their own code. Like so many things with agents, you can start by asking your agent. Google's getting-started guide says to clone Mantis, then ask a coding agent for help: Use the Mantis framework at path/to/mantis to review the code at path/to/your/code. Help me set it up and run the review.Replace both paths with your own. I love how easy that is to try. What to do next: Run Mantis in an isolated environment against a codebase you can safely test. See what it finds, and, if you can, have a security expert check each finding before acting on it. That's it! Which of these grabbed your attention this week, and what will you do about it? Reply and let me know. I read every message. –KM
|
A weekly briefing for business leaders who want the few AI developments that matter.