
Same task, different total bill. Editorial metaphor.
An expense report can cost more to finish even when the model charges less per request. Wrong totals and repeated corrections are part of the bill.
That is the useful tension behind this week's model releases. Elsewhere, Atlas gets a hand built around four fingers, video gains a preview-and-enhance workflow, and biology research puts AI predictions in front of a lab bench.
In Brief: GPT-6.1 Sol and Claude Sonnet 5.5 share standard API rates of $2 per million input tokens and $10 per million output tokens. Their total task costs and benchmark positions differ with reasoning settings and workload. Choosing between them means checking correctness, effort, latency, caching and retries, rather than treating equal token prices as equal value.
Key takeaways
Sol leads Sonnet on Artificial Analysis's aggregate index at high effort; Sonnet leads at maximum effort. The setting changes the comparison.
A robotics audit found that correcting a simulated fan's mass reversed two methods' ranking. Physical conditions belong beside success scores.
Runway's new 480p draft mode offers a practical video workflow: choose the useful shot before enhancing it, then count credits per accepted result.
Featured — What does GPT-6.1 Sol vs Claude Sonnet 5.5 actually cost?

Dated benchmark comparison; effort labels do not mean equal compute budgets.
GPT-6.1 Sol arrived on September 29 in the API, Codex and ChatGPT Work. OpenAI reports stronger coding and computer-use performance than GPT-6 Sol. Claude Sonnet 5.5, announced the day before, keeps Sonnet's standard token prices while Anthropic reports completing tasks with fewer tokens. Anthropic attributes its claimed task savings to efficiency.
For Sol, the quoted standard rates apply to prompts of up to 272,000 input tokens. Long-prompt charges and caching need their own check: cached input rates are $0.10 per million tokens for Sol and $0.20 for Sonnet. Even below that threshold, a chat subscription price and an API task bill are different things.
The Artificial Analysis comparison gives the tradeoff some numbers. In its October 9 Intelligence Index v4.3.2 snapshot, Sol scores 50 against Sonnet's 47 at high effort. At maximum effort, Sonnet scores 56 against Sol's 52. Sonnet was evaluated with default fallback enabled; identical effort labels do not guarantee equal computation.
The weighted average cost per benchmark task also diverges: $0.32 for Sol versus $0.88 for Sonnet at high, and $0.72 versus $5.46 at maximum. These are averages across that evaluation suite, not quotes for your next spreadsheet or bug fix. Its mix of public and private tasks is useful evidence, but cannot represent every job.
So which should you use? Start with the work you can judge.
Imagine a monthly expense sheet with duplicate transactions, inconsistent dates and a few missing categories. Give both models the same copy, the same instructions and the same spending ceiling. Keep a small answer key: the duplicate rows, the correct totals and examples of dates that must survive conversion. A polished explanation cannot compensate for a wrong total.
Record the first result, then include every correction in the bill. If one model costs less initially but requires three follow-ups, the first-request price tells only part of the story. Your review time matters too: a ten-minute manual cleanup may outweigh a few cents of API savings. We covered this broader distinction in The Price of Intelligence Just Changed.
Try a moderate setting for routine work and raise effort when a specific failure justifies it. Repeat across several tasks and track correct results within your budget.
The Signal — The fan that changed a robot leaderboard

Conceptual simulation illustration. Mass and other physics settings changed.
A simulated desk fan weighed just 10 grams. Researchers auditing robot benchmarks adjusted the physical parameters, including changing the fan to 600 grams, then reran the task with the same checkpoints and success scoring. The target pad's mass and inertia also changed.
In their benchmark audit, the baseline's success rate moved from 40% to 38%. DP-Cache-Fast fell from 61% to 33%. The apparent winner changed when the object became harder to move.
This is an author-reported simulation result, not a physical robot trial. Its lesson travels well: before trusting a score, ask what stayed fixed. In a model comparison that includes effort and tools; in robotics it also includes mass, friction and the success check.
AI Briefing — A job that continues after the chat

A proposed review boundary, not a universal permission rule.
OpenAI dots run on cloud computers and use GPT-6 Astra for ongoing work across connected apps. Access is rolling out to eligible plans and markets, so check your account before building a workflow around them.
A useful first assignment might be: monitor three competitors' public release notes, prepare a weekly comparison, and flag changes that could affect our support documentation. That describes a result you can inspect. “Handle our competitive strategy” does not.
Proactive research uses read-only app tools; delegated tasks and action rules have different permissions. Broader teams of dots remain a future direction. The announcement supplies no independent reliability rate.
My suggested boundary is concrete: gathering and drafting can continue, while sending messages or changing customer-facing material waits for review. Persistent work needs equally persistent limits.
DeepSeek's new release is software for a different chip

Conceptual software-porting metaphor, not a performance comparison.
On September 30, DeepSeek released DeepGEMM-Ascend for Huawei Ascend 950 hardware, alongside Ascend support in related attention and kernel tools.
Think of a kernel as a carefully tuned piece of software for a repeated calculation. Fast chips still need those calculations to run efficiently. Porting the software helps developers use another hardware platform without rebuilding every layer around it.
This is open infrastructure, separate from the earlier DeepSeek V4.1 model release. Selected kernel measurements do not establish that an entire Ascend system is faster or cheaper than an NVIDIA system. For readers without Ascend hardware, the significance is the expanding software ecosystem rather than a new model to download tonight.
Robotics Radar — Why Atlas has four fingers

Conceptual four-digit grip; not an exact Atlas replica.
Try holding a drill while pressing its trigger. Gripping the handle is only half the job: one finger needs to move without destabilizing the tool.
Boston Dynamics' new Atlas hand has four fingers and 13 degrees of freedom, up from seven in the previous design. That gives the hand more controlled ways to move. Direct actuation and pressure sensing in the fingers and palm help the robot move objects and detect contact.
The company is balancing dexterity against strength, impact resistance, manufacturing cost and repairability. A fifth finger would add complexity; copying human anatomy is only useful if it improves the jobs the robot needs to perform.
Object reorientation—turning something within the hand—offers a clearer test than a dramatic full-body pose. Can it adjust a part for assembly or keep a tool secure while operating it?
Boston Dynamics shows early work transferring manipulation learned in simulation to hardware. Factory uptime, repeatable recovery from slips and performance across unfamiliar tools still need measurement. Those are the measurements worth watching next. Our robot-surgery simulation explainer examined the same gap between a training environment and real-world validation.
Tool to Try — Preview the shot before paying to enhance it

Proposed exercise illustrated; not actual Runway test outputs.
Runway's October 2 update adds Seedance 2.5 Draft mode in Agent: generate at 480p, then enhance the output you want to keep. The changelog lists it for all plans; that does not mean generations are free or unlimited.
Runway's help guide still lists draft mode as Tools-only, so its instructions lag the Agent announcement. Check which controls your account exposes. Enhancement is a separate charge and the guide specifies a seven-day window.
Here is a proposed exercise; we have not tested the feature. Start with one narrow brief: a five-second shot of steam rising from a cup beside a book, with a fixed camera. Create a few drafts. Decide which has credible steam, stable objects and the timing you need before enhancing it.
The draft stage is useful because many rejected clips fail for reasons resolution will not repair. A wandering cup or an unwanted camera move remains the wrong shot at higher quality.
Keep a simple count: drafts generated, enhanced clips, accepted clips and total credits. Divide credits by accepted clips. Include enhancement failures and any extra retries. That is the number to compare with your previous workflow. Compare this with generating every attempt at full resolution. Drafts cost the same as standard 480p clips; enhancement adds another charge. Whether the workflow saves money depends on how many drafts you reject.
What Comes Next — Biology predictions meet the experiment

Predictions select candidates. Experiments establish evidence.
Biology research rarely lacks possible experiments. Choosing the next informative one is harder.
Microsoft Research's Quine combines a model that connects different kinds of biological data with reasoning and tools. Microsoft describes work with the Broad Institute that ranked compounds for studies of pancreatic-cancer cell states, followed by laboratory assays that tested the predicted changes.
The interesting loop is straightforward: propose a candidate, run the experiment, compare the result with the prediction, then decide what to test next. A wrong prediction can still teach researchers something if the test distinguishes between competing explanations.
Quine is experimental, with initial access through selected collaborations and fellows. The reported lab results are not evidence of a cancer treatment in patients. Nor does rapid candidate selection compress all the preceding research or subsequent clinical work into a weekend.
Watch whether Quine repeatedly selects informative experiments on new questions, with results other researchers can inspect. That would provide stronger evidence than another promising demonstration.
The Intellica Desk
AI, robotics and emerging technology—made clear for curious minds.