Valliance Logo in White
Valliance Logo in White

ISSUE NO.

ISSUE.

04

·

ISSUE NO.

04

·

The Week in AI

Our weekly read on the AI stories that matter to the people building the future enterprise.

ISSUE NO.

ISSUE.

04

·

ISSUE NO.

04

·

The Week in AI

Our weekly read on the AI stories that matter to the people building the future enterprise.

OpenAI introduces Astra for Law, a GPT-6 Astra based offering for law firms and legal technology companies, pairing the model with legal-specific settings, a large legal search index and workflow tools, alongside privacy, governance and integration features including 26 ecosystem plugins.

Spotted on

OpenAI

Why it matters for you

Although Astra for Law is primarily trained on US court filings, the underlying message is pertinent for practices across the globe. A 54% pass rate is a significant improvement, but it's not sufficient for the courtroom. Our experience tells us that many practitioners are slow to adopt legal AI tools, and perhaps this is part of the reason why. The other headline is Open AI's forward deployed engineer model, and it shouldn't be ignored. Law firms that integrate decades of institutional knowledge and experience into frontier models risk losing their competitive advantage. We believe that firms succeed when they build their own data knowledge warehouse, an ontology, that encodes their experience and ways of working. This digitisation of the firm's 'alpha' realises the benefits from the advances in AI without ceding differentiated expertise to tomorrow's competition.

Find out more about:

Anthropic’s new Opus 5.5 undercuts its own previous generation on price, including a 40% lower typical run cost than Opus 5 and faster output generation. It also claims meaningful gains in coding, knowledge work and communication quality, along with strengthened safeguards across safety, cybersecurity, biology and distillation.

Spotted on

Anthropic

Why it matters for you

I’d ask your team one question before anyone signs off on which model is running in production. Do you actually know when it’s you making the decision to change model, versus the vendor making it for you and not saying so? Within the release are two governance details that matter more than the benchmark table. For cybersecurity tasks specifically, Opus 5.5 quietly swaps itself for an older model on certain task types. If you’re running evals pinned to a specific model version, that’s a result you’re no longer able to fully trust without checking. We’re seeing new model releases frequently, with providers continually iterating to keep pace with one another. That makes it increasingly important to have a fully automated evaluation suite in place so that model changes can be assessed, validated, and turned around quickly. A system built around a pinned model version will require active maintenance whenever a provider updates or retires that version, so the ability to rapidly re-run evaluations and confidently migrate becomes critical.

TypeSafe AI's Jev is a System One model: you give it unstructured state (text or JSON) and a set of questions with predefined answers, and it returns typed answers with calibrated probabilities. Brian demoed it live, and it drew a lot of technical questions from the room. Rather than generating text token by token, it evaluates every question in one parallel pass, so you can ask dozens of things about the same state at once with little added latency. Outputs always match the schema, so there's no parsing or retrying. TypeSafe claims sub-second responses at a fraction of LLM cost, with comparable accuracy on these tasks.

Spotted on

TypeSafe

Why it matters for you

The effort moves from prompting and validating LLM output to defining the decision upfront, which is closer to conventional software design and could change how that work is scoped and billed. The obvious fit is classify, route and score steps where an LLM is too slow or unreliable, especially when you need many judgements about the same record. It won't replace anything that needs generated text.

A well-argued piece on turning a personal knowledge habit into something an entire company can rely on, a shared, constantly updated source of company knowledge that both people and AI agents can read from and write to safely. The room’s verdict was that the direction is right, but moving from personal knowledge to an enterprise “company brain” needs much more consideration when it comes to the technology landscape and the potential challenges.

Spotted on

Simon Späti

Why it matters for you

The bigger opportunity here is real. A shared, trustworthy source of company knowledge that both your people and your agents can rely on is worth building toward. What I’d caution against is treating any single vendor’s pattern as the finished answer. The hard parts at enterprise scale are shared definitions across departments, fine-tuned permission controls, and review that doesn’t overload your domain experts. This space is moving fast, and the tooling is still catching up to the ambition.

A piece of design fiction imagining a fully autonomous “dark factory,” a system that, once initialised, needs no further human input and eventually sets its own goals. The room read it less as a blueprint and more as a warning, and used it to sharpen the contrast with Valliance’s own approach to human oversight.

Why it matters for you

This was a genuinely fun read internally, mostly because it kept sparking the conversations we're already having around security, governance and ethics. Strip the fiction away and there’s a real design choice underneath it. Do you build systems that need a human in the loop at every step, or systems that, once switched on, you can only influence rather than control? We think the second one is the wrong bet for an enterprise, however efficient it looks on paper. That’s exactly why governance sits at the centre of how we think about building an enterprise OS, designed in from day one rather than patched on after something breaks.

ISSUE NO.

03

_More of our latest thinking

_More of our latest thinking

_More of our latest thinking

_More of our latest thinking

ISSUE NO.

ISSUE.

04

·

ISSUE NO.

04

·

The Week in AI

Our weekly read on the AI stories that matter to the people building the future enterprise.

ISSUE NO.

ISSUE.

04

·

ISSUE NO.

04

·

The Week in AI

Our weekly read on the AI stories that matter to the people building the future enterprise.

OpenAI introduces Astra for Law, a GPT-6 Astra based offering for law firms and legal technology companies, pairing the model with legal-specific settings, a large legal search index and workflow tools, alongside privacy, governance and integration features including 26 ecosystem plugins.

Spotted on

OpenAI

Why it matters for you

Although Astra for Law is primarily trained on US court filings, the underlying message is pertinent for practices across the globe. A 54% pass rate is a significant improvement, but it's not sufficient for the courtroom. Our experience tells us that many practitioners are slow to adopt legal AI tools, and perhaps this is part of the reason why. The other headline is Open AI's forward deployed engineer model, and it shouldn't be ignored. Law firms that integrate decades of institutional knowledge and experience into frontier models risk losing their competitive advantage. We believe that firms succeed when they build their own data knowledge warehouse, an ontology, that encodes their experience and ways of working. This digitisation of the firm's 'alpha' realises the benefits from the advances in AI without ceding differentiated expertise to tomorrow's competition.

Find out more about:

Anthropic’s new Opus 5.5 undercuts its own previous generation on price, including a 40% lower typical run cost than Opus 5 and faster output generation. It also claims meaningful gains in coding, knowledge work and communication quality, along with strengthened safeguards across safety, cybersecurity, biology and distillation.

Spotted on

Anthropic

Why it matters for you

I’d ask your team one question before anyone signs off on which model is running in production. Do you actually know when it’s you making the decision to change model, versus the vendor making it for you and not saying so? Within the release are two governance details that matter more than the benchmark table. For cybersecurity tasks specifically, Opus 5.5 quietly swaps itself for an older model on certain task types. If you’re running evals pinned to a specific model version, that’s a result you’re no longer able to fully trust without checking. We’re seeing new model releases frequently, with providers continually iterating to keep pace with one another. That makes it increasingly important to have a fully automated evaluation suite in place so that model changes can be assessed, validated, and turned around quickly. A system built around a pinned model version will require active maintenance whenever a provider updates or retires that version, so the ability to rapidly re-run evaluations and confidently migrate becomes critical.

TypeSafe AI's Jev is a System One model: you give it unstructured state (text or JSON) and a set of questions with predefined answers, and it returns typed answers with calibrated probabilities. Brian demoed it live, and it drew a lot of technical questions from the room. Rather than generating text token by token, it evaluates every question in one parallel pass, so you can ask dozens of things about the same state at once with little added latency. Outputs always match the schema, so there's no parsing or retrying. TypeSafe claims sub-second responses at a fraction of LLM cost, with comparable accuracy on these tasks.

Spotted on

TypeSafe

Why it matters for you

The effort moves from prompting and validating LLM output to defining the decision upfront, which is closer to conventional software design and could change how that work is scoped and billed. The obvious fit is classify, route and score steps where an LLM is too slow or unreliable, especially when you need many judgements about the same record. It won't replace anything that needs generated text.

A well-argued piece on turning a personal knowledge habit into something an entire company can rely on, a shared, constantly updated source of company knowledge that both people and AI agents can read from and write to safely. The room’s verdict was that the direction is right, but moving from personal knowledge to an enterprise “company brain” needs much more consideration when it comes to the technology landscape and the potential challenges.

Spotted on

Simon Späti

Why it matters for you

The bigger opportunity here is real. A shared, trustworthy source of company knowledge that both your people and your agents can rely on is worth building toward. What I’d caution against is treating any single vendor’s pattern as the finished answer. The hard parts at enterprise scale are shared definitions across departments, fine-tuned permission controls, and review that doesn’t overload your domain experts. This space is moving fast, and the tooling is still catching up to the ambition.

A piece of design fiction imagining a fully autonomous “dark factory,” a system that, once initialised, needs no further human input and eventually sets its own goals. The room read it less as a blueprint and more as a warning, and used it to sharpen the contrast with Valliance’s own approach to human oversight.

Why it matters for you

This was a genuinely fun read internally, mostly because it kept sparking the conversations we're already having around security, governance and ethics. Strip the fiction away and there’s a real design choice underneath it. Do you build systems that need a human in the loop at every step, or systems that, once switched on, you can only influence rather than control? We think the second one is the wrong bet for an enterprise, however efficient it looks on paper. That’s exactly why governance sits at the centre of how we think about building an enterprise OS, designed in from day one rather than patched on after something breaks.

ISSUE NO.

03

_More of our latest thinking

_More of our latest thinking

_More of our latest thinking

_More of our latest thinking