You name what you want. The rest is left to the chef.
A light agent framework for Ruby — about 800 lines on top of RubyLLM. An agent is an object: fields are state, methods are what the model can call, and the methods you declare without a body are written by the model at runtime.
Ordinary methods are what the model can call — there is no tool registry, so adding a tool is
adding a method. Declared methods have no body: the name and prompt are the
specification, the block is the contract. A field is memory: the chat is fresh every call, so
what carries is what the object keeps — context puts it in the next prompt.
class RefundAgent < Omakase::Agent
instructions "You are the refund desk of an online shop."
describe "Every order this customer placed, newest first. An Order has items and total"
def orders_for(email) = Order.where(email:).order(placed_on: :desc)
describe "What the policy says about a topic, such as :damage or :late"
def policy_on(topic) = POLICY.fetch(topic, "Refunds are allowed within 30 days.")
generates :decide, "Decide this refund, and name the policy you applied.", returns: Refund
end
RefundAgent.decide(email: "ada@example.com", complaint: "the mug arrived cracked")
# => #<struct Refund order_id=1, amount=39.9, reason="Damaged goods are refunded in full…">
The model gets one tool. Its code is evaluated on the agent itself, so the agent's methods and
state are the API; what the code printed and returned comes back as the observation.
It answers with finish(value) — computed, not retyped.
rubyorders_for("ada@example.com").each { |o| puts o.total, o.items.inspect }out39.9
[#<Item name: "Stoneware mug", price: 34.0>, #<Item name: "Shipping", price: 5.9>]rubyputs policy_on(:damage)outDamaged goods are refunded in full, including shipping, within 90 days.rubyfinish(Refund.new(order_id: 1, amount: 39.9, reason: "…"))outanswer accepted
A failure comes back with the line that raised. An answer off contract is rejected into the same loop. Ten tool calls and thirty seconds per execution, then it answers with what it has.
A block is a schema the provider must fill, so callers get validated data rather than text to parse. A scalar is the same thing, unwrapped. A Ruby class means the method hands back the object the code built.
generates :summarize # no prompt: the method name is the prompt
generates :score, returns: :integer # => 7
generates :file_ticket, returns: Ticket # => #<struct Ticket id="A-1">
generates :caption, model: "claude-haiku-4-5" # a cheap model for a cheap job
generates :translate, -> { "Into #{@language}." } # a prompt read at call time
generates :analyze, "Analyze the feedback." do
string :sentiment, enum: %w[positive negative neutral mixed]
array :topics, of: :string
end
One call, straight into the schema. No code runs. Classification, extraction, rewriting.
The loop above. Multi-step work, anything that should use the agent's own methods.
An MCP server's tools arrive as methods on the agent. A skill — a SKILL.md directory,
the same front matter Claude Code uses — arrives as one more, and memory adds two:
remember and recall, by meaning. Nothing new to learn: external services,
curated guidance and what the agent picked up join the list your own methods are on.
class DocsAgent < Omakase::Agent
mcp :files,
transport_type: :stdio,
config: {command: "npx", args: ["-y", "@modelcontextprotocol/server-filesystem", Rails.root.to_s]}
skill "skills/commit-style"
memory
generates :subject_for, "Write the commit subject for this change.", returns: :string
end
Every tool the server lists, described with its arguments. Generated code calls one, then feeds the result to your own method in the same expression.
The front matter's description sits in the prompt. The markdown body costs nothing until the code asks for it — which is all loading on demand has to mean.
Some work is a tree, and the tree is usually already in your database — a comment thread, a category tree, a bill of materials. One agent per node folds it from the leaves up, and the base case is a node with no children, so nothing has to invent how deep to go.
class ThreadAgent < ApplicationAgent
instructions "You sum up a discussion for someone who has not read it."
def roll_up
return said if @comment.replies.empty? # a leaf is its own summary
summarise(comment: said, replies: @comment.replies.map { |reply| self.class.new(reply).roll_up })
end
generates :summarise, "Sum up this comment together with the replies it drew.", returns: :string
end
The model is asked only where there is something to fold. A fresh agent per branch is not ceremony either: siblings then share no state, and one object may not re-enter a generation it is already inside — that is refused, because a nested run opens its own chat with its own tool budget, and nothing would bound the spend. Generated code can start a sub-agent the same way.
Agents live in app/agents, reloading is safe, printing is per-thread so Puma is fine,
and every call lands on the notification bus with its tokens and latency.
# config/initializers/omakase.rb
Omakase.configure do |config|
config.anthropic_api_key = Rails.application.credentials.anthropic_api_key
config.request_timeout = 60
config.instrumenter = ActiveSupport::Notifications
end
class TriageJob < ApplicationJob
def perform(ticket) = ticket.update!(SupportAgent.triage(message: ticket.body))
end
Your validations are the contract. A returns: class that answers
valid? is asked before the answer is handed back, so a generation cannot return a
record your own rules reject — and none of those rules has to be in the prompt.
class Post < ApplicationRecord
validates :slug, length: {maximum: 12}
end
generates :write, "Write a post about the topic.", returns: Post
# => finish rejected: Post is invalid: Slug is too long (maximum is 12 characters)
# and it corrects itself inside the same loop, within its call budget
Identity is a row, state is your tables, and the agent is a value — built for one turn and
thrown away. Multi-turn is context reading the history back, so any pod serves
any turn.
rails console is already the harness: SupportAgent.new(conversation).reply
is one real generation, and every tool is an ordinary method you can call with no model in the
room.
One rule. Generated code runs with instance_eval in your process,
where ActiveRecord and ENV live. Untrusted input belongs to
:predict; :code_act is for work you control.
Executor::Subprocess isolates a crash, not File.
An agent takes any object that quacks like a chat, and one ships with the library — so a test
builds the agent and calls the method, with no network and no cassettes. The class-level call a
job makes has no seam to inject through; chat_factory is that seam.
chat = Omakase::FakeChat.new { {"severity" => "high"} }
assert_equal "high", SupportAgent.new(chat:).triage(message: "broken")[:severity]
# or drive the one tool the way a model would
chat = Omakase::FakeChat.new { |fake| fake.run("finish(stock_of(:apple))") }
# test/test_helper.rb — and the whole suite is off the network
Omakase.chat_factory = ->(**) { Omakase::FakeChat.new { raise "an agent asked for a model" } }
The fake records instructions, schema, tools and
tasks, so the prompt is assertable too. Make the factory raise, and a generation you
forgot to stub fails loudly instead of quietly calling a provider from CI — an injected
chat: still wins.
The tool loop, the schema plumbing and the correction turn after an off-contract answer are what these 800 lines are. Providers, keys, models, streaming and tracing stay RubyLLM's, and stay reachable.
Nothing to register, nothing to keep in sync. A tool's description is describe,
one line above the method — not a JSON schema that drifts from the code it describes.
Not here, on purpose: a checkpoint inside a generation, a sandbox, multi-agent orchestration, a token stream. The library stops where your stack already has an answer.
0.4.0 — usable, not finished. Ruby 3.2+, and any provider RubyLLM supports: Anthropic,
OpenAI, Gemini, Bedrock, Ollama, OpenRouter, and the rest.
gem install omakase-agents
# or, in a Gemfile
gem "omakase-agents"