Sakhanda Wire
NVDA $230.86 +1.09% MSFT $512.80 -0.02% GOOGL $338.24 -1.70% META $725.93 +0.10% AMZN $248.23 -0.37%
← До новин

Посібник із програмування TypeSafe AI Jev: типізовані рішення, калібрована впевненість і спекулятивне розгалуження з моделлю System One

У цьому навчальному посібнику ми працюємо з Jev — першою моделлю System One від TypeSafe AI, яка взагалі не генерує текст: ми надсилаємо їй фрагмент стану програми та набір типізованих запитань, а вона повертає варіанти, оцінки й імовірності «так/ні», на основі яких наш код може безпосередньо виконувати розгалуження. Ми встановлюємо офіційний Python SDK, робимо перший виклик, що одночасно використовує всі три примітиви запитань, і розглядаємо, як форма стану визначає, що саме може знати модель. Потім повторно обчислюємо опубліковану статистику впевненості з повернутих імовірностей, вимірюємо переваги об’єднання десяти запитань в один виклик порівняно з десятьма окремими викликами та створюємо шаблони, для яких призначений API: маршрутизацію з урахуванням упевненості, складене оцінювання зі збереженням ваг у коді, типізовані виклики функцій і підрахунок у спосіб, який модель справді може виконати. На завершення ми формуємо структуру для production: моделі відповідей Pydantic, асинхронний клієнт із розподілом викликів через asyncio, політики повторних спроб, типізовані помилки та поточний журнал, що розраховує вартість усього ноутбука.

Копіювати кодСкопійованоСкористайтеся іншим браузером
import os
import sys
import json
import time
import asyncio
import traceback
import subprocess
from getpass import getpass
 
RESULTS = {}
LEDGER = {"calls": 0, "input_tokens": 0, "output_tokens": 0}
USD_PER_MILLION_INPUT_TOKENS = 0.042          # Jev list price; output tokens are free
 
 
def banner(title):
    print("\n" + "=" * 78)
    print(title)
    print("=" * 78)
 
 
def section(name):
    def wrap(fn):
        def run(*a, **kw):
            banner(name)
            try:
                out = fn(*a, **kw)
                RESULTS[name] = out if isinstance(out, str) else "ok"
                return out
            except Exception as e:
                RESULTS[name] = f"SKIPPED / FAILED -> {type(e).__name__}: {e}"
                print(f"\n[!] {name} did not complete: {type(e).__name__}: {e}")
                traceback.print_exc(limit=3)
                return None
        return run
    return wrap
 
 
banner("0. Install the SDK, load the API key, list the models")
subprocess.run([sys.executable, "-m", "pip", "install", "-q", "typesafe-sdk==0.7.0"], check=True)
import typesafe_sdk
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
 
 
def load_api_key():
    key = os.environ.get("TYPESAFE_API_KEY", "").strip()
    if not key:
        try:
            from google.colab import userdata          # Colab: key stored under the Secrets tab
            key = (userdata.get("TYPESAFE_API_KEY") or "").strip()
        except Exception:
            key = ""
    return key or getpass("TypeSafe API key (console.typesafe.ai/keys): ").strip()
 
 
os.environ["TYPESAFE_API_KEY"] = load_api_key()
client = TypeSafeClient()                              # reads TYPESAFE_API_KEY, defaults to jev-latest
 
print(f"  typesafe-sdk {typesafe_sdk.__version__}  |  Python {sys.version.split()[0]}")
print("  models available to this key:")
for m in client.models.list().models:
    print(f"    {m.name:<14s} released {m.release_date}   {m.description}")
 
 
def ask(state, questions, **kw):
    """One System One call, timed, with its tokens added to the running ledger."""
    t0 = time.perf_counter()
    response = client.system_one(state, questions, **kw)
    ms = (time.perf_counter() - t0) * 1e3
    LEDGER["calls"] += 1
    LEDGER["input_tokens"] += response.usage.input_tokens or 0
    LEDGER["output_tokens"] += response.usage.output_tokens or 0
    return response, ms

Ми встановлюємо typesafe-sdk із фіксацією версії, під яку було написано цей ноутбук, і завантажуємо ключ API з environment-змінної, вкладки Secrets у Colab або прихованого запиту, щоб він ніколи не з’являвся в ноутбуці. TypeSafeClient самостійно читає TYPESAFE_API_KEY і за замовчуванням використовує псевдонім jev-latest; перелік моделей показує, які назви та зафіксовані версії доступні цьому ключу. Невелика допоміжна функція ask обгортає system_one, тому кожен виклик у решті ноутбука вимірюється, а його використання токенів потрапляє до журналу, підсумок якого ми отримаємо наприкінці.

Копіювати кодСкопійованоСкористайтеся іншим браузером
TICKET = {
    "ticket": {
        "subject": "Duplicate charge",
        "messages": [
            {"from": "customer", "text": "I was charged twice for order A-104. This is the second time "
                                         "this year. Please refund the duplicate today."},
            {"from": "support", "text": "We are checking the charges."},
        ],
    },
    "order": {"id": "A-104", "charges": [{"amount_usd": 49, "status": "captured"},
                                         {"amount_usd": 49, "status": "captured"}]},
    "refund_policy": "Duplicate charges are eligible for a full refund within 30 days.",
}
 
 
@section("1. Three primitives, one call: Choice, Score, Noul")
def three_primitives():
    response, ms = ask(TICKET, {
        "department": Choice(
            instructions="Which team should handle this ticket",
            criteria={"billing": "Payment, refund or subscription issues",
                      "technical": "Bugs, outages or integration problems",
                      "sales": "Pricing, plans or account upgrades"},
        ),
        "frustration": Score(
            instructions="How frustrated the customer appears in `ticket.messages[0].text`",
            criteria=["Calm, just stating facts", "Frustrated but civil", "Very angry, strong language"],
        ),
        "refund_requested": Noul(instructions="The customer is explicitly asking for a refund"),
        "policy_supports": Noul(instructions="The stated `refund_policy` covers this situation"),
    })
 
    dept = response.choices["department"]
    print(f"  department       -> {dept.choice!r}   confidence {dept.confidence:.3f}")
    print(f"     probabilities    {({k: round(v, 3) for k, v in dept.probabilities.items()})}")
    fr = response.scores["frustration"]
    print(f"  frustration      -> score {fr.score:.3f} on 0..{len(fr.legend) - 1}   confidence {fr.confidence:.3f}")
    for level, text in fr.legend.items():
        print(f"     {level}: p={fr.probabilities[level]:.3f}  {text}")
    print(f"  refund_requested -> noul {response.nouls['refund_requested'].noul:.3f}")
    print(f"  policy_supports  -> noul {response.nouls['policy_supports'].noul:.3f}")
    print(f"\n  answered by {response.model} in {ms:.0f} ms   "
          f"input tokens {response.usage.input_tokens}, output tokens {response.usage.output_tokens}")
    return f"{dept.choice}, frustration {fr.score:.2f}, refund {response.nouls['refund_requested'].noul:.2f}"
 
 
three_primitives()

Запит System One складається з двох частин: стану, яким може бути будь-який текст, JSON-об’єкт або масив, що описує ситуацію, і словника іменованих запитань. Choice обирає одну мітку з визначених нами критеріїв і повертає ймовірність для кожної мітки; Score розміщує стан на впорядкованій шкалі та повертає рівень, зважений за ймовірністю, тому результат може опинитися між двома рівнями; Noul повертає єдину ймовірність того, що твердження є істинним. Назви запитань задаємо ми, і вони ніколи не передаються моделі, тому саме інструкції містять повний зміст і можуть посилатися на вкладені поля за допомогою шляхів у зворотних лапках. Усі чотири запитання оцінюються в одному запиті, паралельно та незалежно одне від одного, а відповідь містить зафіксовану версію моделі, яка відповідала, і кількість оплачених токенів.

Копіювати кодСкопійованоСкористайтеся іншим браузером
@section("2. State is program state: the same question over a string and over named fields")
def state_shapes():
    question = {"eligible": Noul(
        instructions="The customer is eligible for a refund under the company's written policy",
        criteria={"true": "A policy is present and it covers the customer's situation",
                  "false": "No policy is given, or the policy does not cover the situation"},
    )}
    bare = "I was charged twice for order A-104. Please refund the duplicate."
    as_list = [m["text"] for m in TICKET["ticket"]["messages"]]
    shapes = [("string: the message only", bare),
              ("array : the conversation", as_list),
              ("object: ticket + order + policy", TICKET)]
    print(f"  {'state shape':<34s} {'noul':>6s}   input tokens      ms")
    seen = {}
    for label, state in shapes:
        response, ms = ask(state, question)
        seen[label] = response.nouls["eligible"].noul
        print(f"  {label:<34s} {seen[label]:6.3f}   {response.usage.input_tokens:12d}   {ms:5.0f}")
    print("\n  Only the object carries the policy and the two captured charges; the question")
    print("  is identical in all three calls, so any movement comes from the state.")
    return "noul by state shape: " + ", ".join(f"{v:.2f}" for v in seen.values())
 
 
state_shapes()

Стан — це єдине, що знає модель, тому ми ставимо одне запитання — чи має клієнт право на повернення коштів відповідно до письмової політики компанії — для трьох форм стану. Простий рядок містить лише скаргу; масив додає розмову; JSON-об’єкт додає замовлення з двома списаними платежами та саму політику повернення. Запитання не змінюється, тому будь-яка різниця в повернутій імовірності пояснюється станом, а стовпчик із токенами показує вартість додаткового контексту. Іменовані поля є рекомендованим підходом, коли контекст складається з кількох частин, оскільки інструкції можуть посилатися на них за назвою.

Копіювати кодСкопійованоСкористайтеся іншим браузером
def confidence_from(probabilities):
    """TypeSafe's published statistic: (count x peak - 1) / (count - 1)."""
    p = list(probabilities.values())
    return (len(p) * max(p) - 1) / (len(p) - 1)
 
 
@section("3. Confidence is a statistic of the distribution, and you can recompute it")
def confidence_math():
    tone = Choice(instructions="What is the tone of the message",
                  criteria={"angry": "Upset or hostile", "calm": "Neutral or polite", "excited": "Enthusiastic or eager"})
    urgency = Score(instructions="How soon this needs attention",
                    criteria=["Can wait", "Needs attention this week", "Needs attention today"])
    messages = {
        "clear    ": "This is the third outage this week and nobody answers. Fix it NOW or I cancel today.",
        "ambiguous": "Well. That was certainly an experience. Let me know when you get a chance.",
    }
    print(f"  {'message':<10s} {'choice':<8s} {'API conf':>8s} {'recomputed':>11s}   "
          f"{'score':>6s} {'sum(level*p)':>13s} {'API conf':>9s}")
    worst = 1.0
    for label, text in messages.items():
        response, _ = ask(text, {"tone": tone, "urgency": urgency})
        t, u = response.choices["tone"], response.scores["urgency"]
        expected = sum(level * p for level, p in u.probabilities.items())
        print(f"  {label:<10s} {t.choice:<8s} {t.confidence:8.3f} {confidence_from(t.probabilities):11.3f}   "
              f"{u.score:6.3f} {expected:13.3f} {u.confidence:9.3f}")
        worst = min(worst, t.confidence)
    print("\n  A Noul has no confidence field: its value already is the probability of yes,")
    print("  so 0.5 means undecided, not medium.")
    return f"lowest tone confidence {worst:.2f}"
 
 
confidence_math()

TypeSafe визначає впевненість як статистику, обчислену з розподілу, який уже містить відповідь: кількість варіантів, помножена на максимальну ймовірність, мінус одиниця, поділена на кількість варіантів мінус одиниця. Ми повторно обчислюємо її з імовірностей Choice і порівнюємо з полем confidence, а також повторно обчислюємо Score як суму кожного рівня, помноженого на його ймовірність. Пропускаючи через однакові два запитання різке повідомлення та навмисно нечітке, ми бачимо, як розподіл і, відповідно, впевненість реагують на неоднозначність. Noul взагалі не має поля confidence, оскільки його значення вже є ймовірністю відповіді «так», а значення близько 0,5 означає невизначеність, а не середній рівень.

Копіювати кодСкопійованоСкористайтеся іншим браузером
POSTMORTEM = """Incident 2291 - checkout latency, 14 March. At 09:12 UTC the payments gateway began timing out
for roughly 18 percent of checkout requests in the EU region. The on-call engineer was paged at 09:15 and
acknowledged at 09:21. Initial suspicion fell on the new fraud-scoring service deployed the previous evening,
and it was rolled back at 09:40 with no improvement. At 10:05 the database team found that a connection pool
limit had been lowered from 400 to 40 by an automated configuration sync, which had silently overwritten a
manual override. The limit was restored at 10:11 and error rates returned to baseline by 10:19. Customer
impact: 3,420 failed checkouts and an estimated 61,000 USD in delayed revenue; no data was lost and no
customer data was exposed. Customers were not notified during the incident; the status page was updated at
10:30, after recovery. Follow-ups: alert on pool saturation, require review for configuration-sync overrides,
and add the status page update to the first fifteen minutes of the on-call checklist."""
 
FANOUT = {
    "root_cause": Choice(instructions="What was the root cause of the incident",
                         criteria={"bad_deploy": "A faulty code or service deployment",
                                   "config_change": "An incorrect configuration value",
                                   "capacity": "Organic traffic exceeded provisioned capacity",
                                   "third_party": "A failure at an external vendor",
                                   "unknown": "The text does not establish a cause"}),
    "detected_by": Choice(instructions="How the incident was first detected",
                          criteria={"alerting": "Automated monitoring or paging", "customer": "Customer reports",
                                    "employee": "An employee noticed by chance", "unclear": "Not stated"}),
    "severity": Score(instructions="Severity of customer impact",
                      criteria=["No customer-visible impact", "Minor degradation for a few customers",
                                "A core flow failed for a meaningful share of customers",
                                "Full outage of a core flow for most customers"]),
    "comms_quality": Score(instructions="Quality of customer communication during the incident",
                           criteria=["Customers were informed promptly while it was happening",
                                     "Customers were informed, but late",
                                     "Customers were only informed after recovery, or never"]),
    "data_exposed": Noul(instructions="Customer data was exposed or leaked"),
    "rollback_helped": Noul(instructions="Rolling back the fraud-scoring service resolved the incident"),
    "human_error": Noul(instructions="A person making a manual mistake directly caused the incident"),
    "has_followups": Noul(instructions="The text lists concrete follow-up actions"),
    "revenue_lost": Noul(instructions="Revenue was permanently lost, as opposed to delayed"),
    "eu_only": Noul(instructions="The impact was limited to the EU region"),
}
 
 
def value_of(answer):
    for field in ("choice", "score", "noul"):              # a score of 0.0 is a real value, not a miss
        if hasattr(answer, field):
            return getattr(answer, field)
 
 
@section("4. Speculative fan-out: ten questions in one call versus ten calls")
def fan_out():
    batched, batched_ms = ask({"postmortem": POSTMORTEM}, FANOUT)
    batched_tokens = batched.usage.input_tokens
    seq_ms, seq_tokens, agree = 0.0, 0, 0
    print(f"  {'question':<16s} {'one call':>10s} {'own call':>10s}")
    for name, q in FANOUT.items():
        single, ms = ask({"postmortem": POSTMORTEM}, {name: q})
        seq_ms, seq_tokens = seq_ms + ms, seq_tokens + single.usage.input_tokens
        a, b = value_of(batched.answers[name]), value_of(single.answers[name])
        same = a == b if isinstance(a, str) else abs(a - b) < 0.05
        agree += same
        fmt = (lambda v: f"{v:>10s}") if isinstance(a, str) else (lambda v: f"{v:10.3f}")
        print(f"  {name:<16s} {fmt(a)} {fmt(b)}   {'same' if same else 'differs'}")
    print(f"\n  one call : {batched_ms:7.0f} ms   {batched_tokens:6d} input tokens")
    print(f"  ten calls: {seq_ms:7.0f} ms   {seq_tokens:6d} input tokens")
    print(f"  -> {seq_ms / batched_ms:.1f}x faster and {seq_tokens / batched_tokens:.1f}x fewer tokens; "
          f"{agree}/{len(FANOUT)} answers agree, because questions never see each other")
    return f"{seq_ms / batched_ms:.1f}x faster, {seq_tokens / batched_tokens:.1f}x cheaper, {agree}/{len(FANOUT)} agree"
 
 
fan_out()

Оскільки запитання в межах одного запиту не можуть бачити відповіді одне одного, ми можемо заздалегідь поставити всі потенційно потрібні запитання, зокрема ті, що мають значення лише для однієї гілки, а потім прочитати тільки релевантні відповіді. Ми об’єднуємо десять запитань про звіт щодо інциденту — два Choice, два Score і шість Noul — в один виклик, потім повторюємо кожне з них в окремому виклику та порівнюємо час виконання, кількість вхідних токенів і відповіді. Стан надсилається один раз, а не десять, завдяки чому скорочуються і затримка, і кількість токенів. Стовпчик узгодженості безпосередньо перевіряє твердження про незалежність: запитання має отримувати однакову відповідь незалежно від того, надсилається воно разом з іншими чи окремо.

Копіювати кодСкопійованоСкористайтеся іншим браузером
INTENT = Choice(
    instructions="What the user wants the banking assistant to do",
    criteria={"check_balance": "See a balance or recent transactions",
              "approve_transfer": "Send or approve a transfer of money",
              "dispute_charge": "Contest a charge they do not recognise",
              "close_account": "Close the account permanently",
              "other": "Anything else, or not clear enough to act on"},
)
STAKES = {"check_balance": 0.50, "dispute_charge": 0.70, "approve_transfer": 0.85, "close_account": 0.90}
 
 
def route(answer):
    if answer.choice == "other" or answer.confidence < 0.50:
        return "-> human"
    bar = STAKES[answer.choice]
    return f"-> run {answer.choice}" if answer.confidence >= bar else f"-> confirm first (needs {bar:.2f})"
 
 
@section("5. Confidence-gated routing: the bar rises with the stakes")
def gated_routing():
    inbox = ["how much is in my checking account",
             "send 2,000 to my landlord like last month",
             "i guess maybe move some money around? not sure",
             "there's a 89.99 charge from a gym i never joined",
             "shut everything down, i'm done with this bank",
             "what's the weather like in lisbon"]
    print(f"  {'message':<50s} {'intent':<17s} {'conf':>5s}  decision")
    acted = 0
    for text in inbox:
        response, _ = ask(text, {"intent": INTENT})
        a = response.choices["intent"]
        decision = route(a)
        acted += decision.startswith("-> run")
        print(f"  {text[:50]:<50s} {a.choice:<17s} {a.confidence:5.2f}  {decision}")
    print(f"\n  thresholds live in code: {STAKES}")
    return f"{acted}/{len(inbox)} messages acted on automatically"
 
 
gated_routing()

Типізовані відповіді мають значення лише тоді, коли код навколо них визначає, наскільки високої впевненості потребує дія. Ми класифікуємо кожне повідомлення за наміром і здійснюємо маршрутизацію на основі двох чинників: самого наміру та того, чи перевищує його впевненість поріг, що зростає разом із ризиками — від 0,5 для перегляду балансу до 0,9 для закриття рахунку. Усе, що класифіковано як other або має впевненість нижче 0,5, передається людині; розпізнаний намір, що не досягає свого порога, спочатку підтверджується користувачем. Пороги є звичайними значеннями Python, тому допустимий рівень ризику можна переглядати, версіонувати й тестувати так само, як будь-який інший код, а не приховувати в промпті.

Копіювати кодСкопійованоСкористайтеся іншим браузером
DIMENSIONS = {
    "python_depth": Score(instructions="Depth of hands-on Python engineering experience", criteria=[
        "No Python mentioned", "Scripts or notebooks only", "Ships production Python services",
        "Designs Python libraries or frameworks used by others"]),
    "ml_systems": Score(instructions="Experience running machine learning systems in production", criteria=[
        "None mentioned", "Trained models offline only", "Deployed and monitored models in production",
        "Owned large-scale training or serving infrastructure"]),
    "leadership": Score(instructions="Evidence of leading people or projects", criteria=[
        "None mentioned", "Mentored individuals", "Led a project or a small team",
        "Managed several teams or an organisation"]),
    "communication": Score(instructions="Evidence of clear written or public communication", criteria=[
        "None mentioned", "Internal docs only", "Public posts or talks", "Widely read writing or major conference talks"]),
}
CANDIDATES = {
    "Asha":  "Eight years of Python; maintains an open-source data validation library with 4k stars. "
             "Deployed fraud models at a bank and ran their monitoring. Mentors two juniors. Writes a technical blog.",
    "Bruno": "Engineering manager for three teams (22 people). Wrote Java for a decade, some Python scripting. "
             "Sponsored the company's ML platform but did not build it. Keynoted two industry conferences.",
    "Chen":  "PhD in statistics; trains models in notebooks, no production deployments. Python for analysis. "
             "Teaching assistant for two courses. Several internal reports.",
    "Dara":  "Built and owned the serving infrastructure for a recommender at 40k requests per second in Python "
             "and C++. Led a five-person platform team. Internal design docs only.",
}
WEIGHTS = {"senior IC": {"python_depth": .40, "ml_systems": .40, "leadership": .05, "communication": .15},
           "team lead": {"python_depth": .15, "ml_systems": .25, "leadership": .45, "communication": .15}}
 
 
@section("6. Composite scoring: atomic judgments from the model, weights from code")
def composite_scoring():
    table = {}
    for name, bio in CANDIDATES.items():
        response, _ = ask({"candidate_bio": bio}, DIMENSIONS)
        table[name] = {d: response.scores[d].score / (len(q.criteria) - 1) for d, q in DIMENSIONS.items()}
    print(f"  {'':<7s}" + "".join(f"{d:>15s}" for d in DIMENSIONS) + "   (each normalised to 0..1)")
    for name, row in table.items():
        print(f"  {name:<7s}" + "".join(f"{row[d]:15.2f}" for d in DIMENSIONS))
    winners = {}
    for role, w in WEIGHTS.items():
        ranked = sorted(table, key=lambda n: -sum(w[d] * table[n][d] for d in w))
        winners[role] = ranked[0]
        print(f"\n  ranking for {role:<10s}: " +
              "  >  ".join(f"{n} {sum(w[d] * table[n][d] for d in w):.2f}" for n in ranked))
    print("\n  Two rankings, four model calls: changing the weights re-ran no inference.")
    return ", ".join(f"{role}: {who}" for role, who in winners.items())
 
 
composite_scoring()

Складене оцінювання звужує завдання моделі, а політику робить явною. Для кожного кандидата ми ставимо чотири запитання Score, кожне з яких описує конкретні ситуації, а не ступені прояву ознаки, нормалізуємо кожну оцінку за її максимальним рівнем і зберігаємо отриману таблицю. Далі рейтинг визначається звичайною арифметикою: один набір ваг — для старшого індивідуального спеціаліста, інший — для керівника команди. Оскільки оцінки зберігаються окремо від ваг, зміна пріоритетів миттєво змінює рейтинг кандидатів і не потребує нового висновку моделі. Кожну позицію в рейтингу можна простежити до виміру, який її сформував.

Копіювати кодСкопійованоСкористайтеся іншим браузером
ROOMS = {"living_room": None, "bedroom": None, "kitchen": None, "office": None}
 
 
def set_lights(room, state):
    return f"lights in {room} -> {state}"
 
 
def set_thermostat(room, mode):
    return f"thermostat in {room} -> {mode}"
 
 
def play_music(room, genre):
    return f"playing {genre} in {room}"
 
 
TOOLS = {"set_lights": (set_lights, "state"), "set_thermostat": (set_thermostat, "mode"),
         "play_music": (play_music, "genre")}
CALL_SPEC = {
    "tool": Choice(instructions="Which smart-home function the command asks for",
                   criteria={"set_lights": "Turn lights on, off, or dim them",
                             "set_thermostat": "Make a room warmer, cooler, or set eco mode",
                             "play_music": "Play music or audio",
                             "none": "Not a smart-home command this system supports"}),
    "room": Choice(instructions="Which room the command refers to", criteria=ROOMS),
    "state": Choice(instructions="If this is a lights command: the requested light state",
                    criteria={"on": None, "off": None, "dim": None}),
    "mode": Choice(instructions="If this is a thermostat command: the requested mode",
                   criteria={"heat": "Warmer", "cool": "Cooler", "eco": "Energy saving"}),
    "genre": Choice(instructions="If this is a music command: the requested genre",
                    criteria={"jazz": None, "classical": None, "rock": None, "ambient": None}),
}
 
 
@section("7. Typed function calling, and counting the way Jev can do it")
def function_calling():
    commands = ["it's freezing in the office, warm it up", "kill the lights in the bedroom",
                "put on something mellow and jazzy in the kitchen", "order me a pizza"]
    dispatched = 0
    for text in commands:
        response, ms = ask(text, CALL_SPEC)                # every argument asked speculatively, one call
        c = response.choices
        tool = c["tool"].choice
        if tool == "none":
            print(f"  {text!r:<52s} -> no tool (confidence {c['tool'].confidence:.2f})")
            continue
        fn, arg = TOOLS[tool]
        weakest = min(c["tool"].confidence, c["room"].confidence, c[arg].confidence)
        print(f"  {text!r:<52s} -> {tool}(room={c['room'].choice!r}, {arg}={c[arg].choice!r})  "
              f"weakest judgment {weakest:.2f}, {ms:.0f} ms")
        print(f"  {'':<52s}    {fn(c['room'].choice, c[arg].choice)}")
        dispatched += 1
 
    basket = ["mango", "spanner", "kiwi", "router", "plum", "stapler", "fig", "lychee"]
    response, _ = ask({"items": basket},
                      {f"item_{i}": Noul(instructions=f"`items[{i}]` is the name of a fruit") for i in range(len(basket))})
    probs = [response.nouls[f"item_{i}"].noul for i in range(len(basket))]
    print("\n  counting: one Noul per item, summed in code (Jev does not count reliably in one question)")
    print("  " + "  ".join(f"{item}={p:.2f}" for item, p in zip(basket, probs)))
    count = sum(p > 0.5 for p in probs)
    print(f"  fruits counted: {count} of {len(basket)}")
    return f"{dispatched}/{len(commands)} commands dispatched from typed answers; counted {count} fruits"
 
 
function_calling()

Виклик функцій перетворюється на набір запитань із фіксованою множиною відповідей: один Choice обирає інструмент, включно з явним варіантом none для непідтримуваних команд, а по одному Choice для кожного аргументу запитується спекулятивно в тому самому запиті. Код читає лише аргументи, що належать вибраному інструменту, повідомляє найслабше судження як рівень упевненості всього виклику, а потім виконує звичайну функцію Python із перевіреними значеннями з переліченої множини. Друга частина використовує документований обхідний шлях: Jev ненадійно рахує в межах одного запитання, тому ми ставимо по одному Noul для кожного елемента в одному запиті, а суму обчислюємо в коді.

Копіювати кодСкопійованоСкористайтеся іншим браузером
import concurrent.futures
from typesafe_sdk import (AsyncTypeSafeClient, ChoiceAnswer, NoulAnswer, RetryPolicy, ScoreAnswer,
                          SystemOneResponse, TypeSafeAPIError, TypeSafeError)
 
 
class TicketDecision(SystemOneResponse):
    """Declare the answers you expect and read them as attributes, validated by Pydantic."""
    department: ChoiceAnswer
    frustration: ScoreAnswer
    refund_requested: NoulAnswer
 
 
TRIAGE = {
    "department": Choice(instructions="Which team should handle this ticket",
                         criteria={"billing": "Payment, refund or subscription issues",
                                   "technical": "Bugs, outages or integration problems",
                                   "sales": "Pricing, plans or account upgrades"}),
    "frustration": Score(instructions="How frustrated the customer appears",
                         criteria=["Calm, just stating facts", "Frustrated but civil", "Very angry, strong language"]),
    "refund_requested": Noul(instructions="The customer is explicitly asking for a refund"),
}
QUEUE = ["My invoice shows two seats but I only have one user.", "The export button does nothing in Safari.",
         "Can I get a discount if I pay annually?", "Your API returns 500 on every request since this morning!!",
         "I want my money back for last month, the product never worked.", "How do I add a teammate?",
         "Webhooks stopped firing after your update.", "Do you offer a plan for nonprofits?",
         "Charged after I cancelled. Refund this immediately.", "The dashboard is slow but usable.",
         "Is there an on-prem version?", "Login emails never arrive."]
 
 
def run_async(coro):
    """Works in a plain script and inside Jupyter/Colab, where an event loop is already running."""
    try:
        asyncio.get_running_loop()
    except RuntimeError:
        return asyncio.run(coro)
    with concurrent.futures.ThreadPoolExecutor(max_workers=1) as pool:
        return pool.submit(asyncio.run, coro).result()
 
 
async def triage_all(tickets):
    retry = RetryPolicy(max_retries=3, backoff_initial=0.5, backoff_max=4.0, timeout=20.0)
    async with AsyncTypeSafeClient(retry=retry, timeout=10.0) as aclient:
        t0 = time.perf_counter()
        results = await asyncio.gather(*(aclient.system_one(t, TRIAGE, response_model=TicketDecision)
                                         for t in tickets))
        return results, (time.perf_counter() - t0) * 1e3
 
 
@section("8. Production shape: typed response models, async fan-out, retries, errors")
def production():
    results, wall_ms = run_async(triage_all(QUEUE))
    for r in results:
        LEDGER["calls"] += 1
        LEDGER["input_tokens"] += r.usage.input_tokens or 0
        LEDGER["output_tokens"] += r.usage.output_tokens or 0
    print(f"  {len(QUEUE)} tickets triaged concurrently in {wall_ms:.0f} ms wall time "
          f"({wall_ms / len(QUEUE):.0f} ms per ticket amortised)\n")
    print(f"  {'ticket':<58s} {'department':<10s} {'frustr.':>7s} {'refund':>7s}")
    for text, r in zip(QUEUE, results):                    # attribute access, no dict lookups, no parsing
        print(f"  {text[:58]:<58s} {r.department.choice:<10s} {r.frustration.score:7.2f} {r.refund_requested.noul:7.2f}")
 
    print("\n  errors are typed too:")
    try:
        client.system_one("anything", {})
    except TypeSafeError as e:
        print(f"    empty questions, caught before any request : {type(e).__name__}: {e}")
    try:
        client.system_one("anything", {"q": Noul(instructions="Is this a test")}, model="jev-does-not-exist",
                          retry=RetryPolicy(max_retries=0))
    except TypeSafeAPIError as e:
        print(f"    unknown model, rejected by the API         : {type(e).__name__} (HTTP {e.status})")
    refunds = sum(r.refund_requested.noul > 0.5 for r in results)
    return f"{len(QUEUE)} tickets in {wall_ms:.0f} ms; {refunds} refund requests flagged"
 
 
production()

Чотири деталі перетворюють приклади на сервіс ухвалення рішень. Наслідування SystemOneResponse і оголошення очікуваних відповідей забезпечують доступ до атрибутів із перевіркою Pydantic, тому типізоване рішення залишається типізованим у всьому застосунку, а не перетворюється на пошук у словнику. Оскільки кожен запит є незалежним, черга заявок є чергою незалежних рішень: AsyncTypeSafeClient разом з asyncio.gather надсилає їх паралельно, а допоміжна функція run_async дає змогу використовувати той самий код у скрипті та ноутбуці, де цикл подій уже запущено. RetryPolicy обмежує кількість повторних спроб, затримку між ними та загальний ліміт часу для кожного виклику. Помилки також типізовані: порожній набір запитань відхиляється ще до надсилання запиту, а невідома назва моделі повертається API як підклас TypeSafeAPIError, що містить HTTP-статус.

Копіювати кодСкопійованоСкористайтеся іншим браузером
banner("SUMMARY")
for name, res in RESULTS.items():
    print(f"  {name:<86s} {res}")
cost = LEDGER["input_tokens"] / 1e6 * USD_PER_MILLION_INPUT_TOKENS
client.close()
print(f"\n  whole tutorial: {LEDGER['calls']} calls, {LEDGER['input_tokens']:,} input tokens, "
      f"{LEDGER['output_tokens']:,} output tokens (free)  ->  about ${cost:.5f}")
print("""
Where to go next
 - Patterns: docs.typesafe.ai/patterns  (fan-out, confidence routing, composite scoring, intent routing)
 - Cookbooks: re-ranking, RAG passage filtering, citation checks, LLM guardrails, hierarchical classification
 - Known rough edges of jev-1.13: docs.typesafe.ai/model-jaggedness/jev-1.13  (literal reading, arithmetic,
   counting, date comparison, large irrelevant state)
 - Compare against an LLM on the same questions: github.com/typesafe-ai/system-one-adapter-python
 - Pin a version for production: TypeSafeClient(model="jev-1.13.0"); response.model reports what answered
""")

У підсумку виводиться однорядковий результат, який повернув кожен розділ, а потім підсумовується журнал, до якого додавався кожен виклик: кількість запитів, вхідні та вихідні токени й вартість за опублікованою ціною вхідних токенів, причому вихідні токени безкоштовні.

Підсумовуючи, ми використовували Jev так, як її задумано: як джерело невеликих типізованих суджень, які комбінує код, а не як генератор тексту, якому потрібно надсилати промпти та розбирати результати. Кожна відповідь надходила у вигляді мітки, рівня або ймовірності з відповідним розподілом, що давало нам змогу задавати пороги, ваги й правила маршрутизації в Python, де їх можна тестувати. Об’єднання запитань над спільним станом скорочує і затримку, і кількість токенів, оскільки стан передається один раз для кожного елемента; Noul замінюють підрахунок, якому не можна довіряти моделі, а Choice із фіксованою множиною варіантів перетворюють команди природною мовою на перевірені виклики функцій. Production-компоненти — типізовані моделі відповідей, асинхронний клієнт, політики повторних спроб і типізовані помилки — є невеликими, а журнал розраховує вартість усього ноутбука. Залишається те, чого жоден SDK не може зробити за нас: перевірити запитання, критерії та пороги на власних даних, перш ніж довірити їм реальні дії.


Перегляньте ПОВНИЙ КОД тут. Уся подяка досліднику цього проєкту. Також підписуйтеся на нас у Twitter і не забудьте приєднатися до нашого ML-сабреддіту з понад 150 тисячами учасників та підписатися на нашу розсилку. Стривайте! Ви користуєтеся Telegram? тепер ви також можете приєднатися до нас у Telegram.

Потрібно співпрацювати з нами для просування вашого GitHub-репозиторію, сторінки Hugging Face, запуску продукту, вебінару тощо? Зв’яжіться з нами

Перекладено автоматично з англійської. Оригінал статті — за посиланням нижче.

Вперше опубліковано виданням MarkTechPost

Читати оригінал на MarkTechPost ↗

Текст і зображення належать MarkTechPost і наводяться тут із зазначенням авторства та посиланням на оригінальну публікацію.

← До новин

Ще новини

Усі останні новини