THE ELLAMI GUIDE
Ellami documentation
Everything you need to work with Ellami: launch Agentic IDE, work with agents, choose models and connect the API. Pick a guide and start building.
Updated 10 October 2026 · Agentic IDE 0.1.10
Start in a few steps
Install Ellami Agentic IDE
Ellami runs on macOS. The public desktop download is being prepared.
Sign in and update
Start signing in through Ellami Console from the app, open the code page and approve access in your account. Return to the IDE after approval. An API key is an alternative sign-in method. To update, quit Ellami, open the new DMG and replace the app in Applications. Check the version in Settings → About.
Connect an API key
- Sign in to Ellami Console and open API keys.
- Create a separate key for each device or app. Set a clear name and the limit you need.
- Copy the secret immediately: the full key appears once. Connect it in the Agentic IDE or store it on your server as ELLAMI_KEY.
Keep keys in your trusted Agentic IDE client or on your server. Do not expose them in website code, public repositories or messages. Revoke an exposed key in Console and create a replacement.
Work with your project
- Task
Goal & context
- Skills
Relevant guidance
- Tools
Files & commands
- Verify
Evidence & result
Skills complement each other. Available tools carry out the actions.
Open a project folder in the Agentic IDE. Your agent uses files, the editor and terminal on your computer. Describe the goal, expected result and constraints: what to change, how to verify it and which parts of the project to preserve.
Fix the sign-in form: show the error next to the field,
preserve the current design and verify the flow in a browser.Choose a model and effort before starting. Review proposed changes and the result. Built-in skills supply guidance for the task; an additional agent can work in parallel and its requests also consume your balance.
Chat, Preview and verification
You can start a chat without a project: the agent works in your home folder and can open a newly created folder in the IDE on request. The right panel contains files, Preview and a browser. Use HTML for static pages; for Vite or Next.js, start the dev server first and enter its local address. Inspect changed files and verification results after the task. File undo does not reverse commands or external actions.
Copilot
The Copilot button opens a learning assistant: it asks about your experience and goal, then explains code and agent messages. Ask Copilot about a message from its menu. Sharing the open file is opt-in. Copilot does not edit files or run commands.
Task limit and voice input
Set a new-task limit in USD in Settings → General; leave it blank to use the server limit. The next request is blocked when the budget is exhausted. A large cost estimate requires separate confirmation. The microphone records speech; stopping sends the transcript together with your draft. A transcription failure requires a new recording.
Continue from your phone
In the IDE, open Settings → Connection and enable phone connections. Open Remote in your phone browser, sign in to the same account and select your Mac. Compare the connection code and approve the phone on your Mac. The computer must stay on and the Agentic IDE connected: files and commands remain on your Mac.
Open Remote ↗From your phone you can send tasks, choose model and effort, approve actions, stop work and read project files. On iPhone, use Share → Add to Home Screen. Preview forwards local pages and GET resources; form submission and HMR/WebSocket connections are unsupported. Notifications require browser permission and, on iPhone, an installed PWA.
Create images and video
Open Graphics in Console and choose Images or Video. Describe the result, optionally add your own reference and select a model. Image quality and variation count are available where supported. For video, choose from the model’s durations, formats and resolutions; one generation creates a single scene.
Download the result in time: images and video are available for five minutes after readiness, then access expires and files are deleted. Reloading the page hides previous-session results. Keep your downloaded copy.
Balance is reserved before generation. A video’s “from” price is a starting floor, not an exact quote for your scene. Unused reserve returns after settlement. Cancelling after provider submission does not guarantee a refund of incurred costs. Check service status in Console if unavailable; video requires a running execution server.
Public API for your application
- Your server
model + messages
Ellami APIYour selected provider’s model
OpenAI
Anthropic
Google
DeepSeek
- JSON response
choices[0].message.content
One Ellami key and request format. Check the HTTP status before processing the response.
A separate OpenAI-compatible text API. It accepts model and messages and returns JSON or an SSE stream. Keep your key on the server.
https://www.ellami.pro/api/v1/chat/completionscurl https://www.ellami.pro/api/v1/chat/completions \
-H "Authorization: Bearer $ELLAMI_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: my-request-001" \
-d '{"model":"gpt-6-luna","reasoning_effort":"high","messages":[{"role":"user","content":"Hello, Ellami"}],"max_tokens":256,"stream":false}'const response = await fetch("https://www.ellami.pro/api/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.ELLAMI_KEY}`,
"Content-Type": "application/json",
"Idempotency-Key": crypto.randomUUID()
},
body: JSON.stringify({
model: "gpt-6-luna",
reasoning_effort: "high",
messages: [{ role: "user", content: "Hello, Ellami" }],
max_tokens: 256,
stream: false
})
});
const result = await response.json();
if (!response.ok) throw new Error(result.message || result.error?.message || "Request failed");
console.log(result.choices[0].message.content);import os, uuid, requests
response = requests.post(
"https://www.ellami.pro/api/v1/chat/completions",
headers={
"Authorization": "Bearer " + os.environ["ELLAMI_KEY"],
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"model": "gpt-6-luna",
"reasoning_effort": "high",
"messages": [{"role": "user", "content": "Hello, Ellami"}],
"max_tokens": 256,
"stream": False,
},
timeout=120,
)
response.raise_for_status()
print(response.json()["choices"][0]["message"]["content"])| Field / header | Purpose |
|---|---|
Authorization | Bearer ELLAMI_KEY |
model | An ID from GET /api/v1/models. |
reasoning_effort | Model effort: low, medium, high, xhigh or max, when reasoning is supported. |
messages | A non-empty array of messages with role and content. |
max_tokens | Maximum output tokens within the model’s limits; examples use 256. |
stream | false — JSON; true — SSE when the model has capabilities.streaming enabled. |
Idempotency-Key | A unique operation ID. Reuse the same key and request body when retrying that operation. |
Response text: choices[0].message.content. Check the HTTP status first. Generate a new Idempotency-Key for a new operation. The examples use GPT-6 Luna: first confirm its id appears in the catalog. Keep the original Idempotency-Key when retrying. Streaming is enabled for all text models authorized for your key with capabilities.streaming: true. Streaming capability does not guarantee current provider availability. Multiple choices with n > 1, Responses and a native Anthropic API are not supported yet.
Pricing and caching
Public API uses the model developers’ official Standard rates with no Ellami markup. /api/v1/models returns USD per million token rates in pricing: input, output, cachedInput and long-context priceTiers. Active promotions apply automatically. The provider manages prompt caching; discounts apply only to verified cached_tokens in usage. Ellami does not cache answers across separate requests. Retrying an Idempotency-Key returns the saved result without another charge.
For OpenAI-compatible clients set baseURL / base_url to https://www.ellami.pro/api/v1 and apiKey / api_key to your Ellami key. Use chat.completions and a model ID from the catalog. For streaming set stream: true and stream_options: { include_usage: true }. Keep the key in a server environment variable. A workspace defaults to 60 API requests per minute and 8 concurrent requests; HTTP 429 includes Retry-After. Errors return an error object with message, type, param and code; X-Request-Id identifies the operation.
Streaming and recovery
Text and function arguments arrive in choices[0].delta. The final finish_reason, ellami_usage and [DONE] are emitted after settlement. HTTP 200 alone does not confirm successful stream completion: a later failure arrives as an SSE error event. After interruption check GET /api/v1/requests/{requestId} with the same key: settled includes the saved response; running needs time; uncertain needs support reconciliation. Retrying the same Idempotency-Key and original body replays the saved stream without another charge. Cancellation stops reading; unknown usage retains the reserve.
Confirming an estimate
HTTP 409 with code: PREFLIGHT_REQUIRED means the estimate in details.estimatedUsd needs approval first. After approval, repeat the original request with the same Idempotency-Key, adding ellami_quote_id from details.quoteId and ellami_confirm_estimate: true. Obtain a fresh estimate if the body changes or the quote expires. Do not approve automatically without the user’s decision.
A successful response contains usage and ellami_usage with usage and charge data. Errors have code and message; inspect them before reading choices. For request_in_progress, keep the original operation key; for reconciliation_required, wait for support reconciliation.
Models and current prices
Key models from the Ellami catalog. Base provider rates in USD are shown below; check Console for the final rate and availability. Catalog checked: 2026-10-07.
Code, documents and complex tasks in a single workflow.
A lightweight model for code, fast responses and high-volume tasks.
Feature development, bug fixing and document work.
Complex code, reviews and long-running agentic tasks.
Multimodal analysis, code and multi-step tasks.
Code, reasoning and visual analysis with a large context.
Code, visual understanding and long-video analysis.
Fast multimodal model for code and visual analysis.
Multimodal development and agentic tasks.
Image generation and editing from prompts and references.
Image generation and editing with output up to 4K.
Fast video generation with audio for iteration and finished content.
The catalog publishes available Public API models and customer prices. Use a canonical model id from the response rather than its display name.
https://www.ellami.pro/api/v1/modelscurl https://www.ellami.pro/api/v1/models \
-H "Authorization: Bearer $ELLAMI_KEY"Agentic IDE cost depends on the model and actual input, output and cached tokens. The Public API has separate input, output and cache rates. Check prices and usage in Console before extended work. Console’s Models page describes a broader catalog; verify a specific Public API id through GET /api/v1/models. An empty list means public routes have not yet been configured.
Balance, limits and bonuses
Usage is charged to your balance at server prices. Funds are reserved before a request; the unused portion returns after the response. A reserve reduces your available balance but is not yet the final charge. IDE, API, agents, and graphics share one USD balance. Console shows your remaining balance and usage history. A key limit restricts that key’s spending; a task budget helps control longer agent work.
Get $1 once when creating a new account with ELLAMI1 and verifying your email.
Discount on your first top-up — once.
On your first top-up you pay 15% less than the selected amount and receive the full amount on your balance. The server grants the welcome balance and discount once. Promo balance cannot be withdrawn or transferred. One-time top-ups start at $3 with no required subscription. Enable automatic top-ups separately when the payment method is available; set a threshold, amount and monthly cap. Check the final amount and payment status in Console; test payments do not charge money.
If something goes wrong
| HTTP | What to check |
|---|---|
| 400 | Request parameters, model and messages; for stream: true check the model’s capabilities.streaming. |
| 401 / 403 | Key, expiration and workspace access. |
| 402 | Balance, key limit and task budget. |
| 409 | Read code: PREFLIGHT_REQUIRED needs estimate approval; quote_expired needs a new quote; idempotency_conflict needs another key for a new operation. Do not retry uncertain paid work with a new key. |
| 429 | Reduce request frequency and retry after a delay. |
| 410 | The Graphics result has expired. Start a new generation if needed. |
| 503 | Check model availability and retry later. |
For Remote, check your Mac and Agentic IDE connection. For key issues, create a new key and reconnect it.
Open console ↗