AI Access#
Open AI Access to onboard people and control who can use which AI capabilities. Invite members, set spend and quota limits, choose allowed providers, and enable cost-optimization settings before users open Chat or call the API.
Use this page when you need to:
- Invite new members or resend pending invitations.
- Cap spend, tokens, or requests for a user or team.
- Restrict which providers a member can use in Chat.
- Turn context compression on or off to manage inference cost (savings appear on Optimizations).
- Group users into teams for shared limits and access scoping.
From this page you can learn who is already onboarded, which invitations are still pending, how limits and allowed providers are assigned, and whether optimization mode is enabled for each principal.
Next: follow the user onboarding workflow, then configure limits, allowed providers, and compression as needed. Register provider connections first under Providers if the organization catalog is empty.
The page has two tabs:
- Users — Invite members, manage invitations, set limits, assign teams and providers, and enable compression. See Users tab.
- Teams — Create teams and apply shared limits, settings, and compression. See Teams tab.
Users tab#
Use the Users tab to onboard members and control per-user access. Open AI Access → Users, then:
- Invite users — click + INVITES, then + INVITE, to send organization invitations.
- Download the user list — click DOWNLOAD and choose XLSX spreadsheet or JSON file.
- Open Invite Management — click INVITES to monitor pending invitations and Resend when needed.
- Configure user limits — in the Limits column, open the pencil control to set budget, tokens, or requests.
- Assign users to teams — in the Team column, open the pencil control and select a team (create teams on the Teams tab first).
- Configure allowed AI providers — in the Allowed providers column, assign the provider connections the user may use in Chat.
- Grant access to MCP servers — assign approved MCP Servers to the user so tools are available in Chat and agent workflows.
- Enable or disable context compression — use the Enable context compression toggle on the user row.
- Manage user settings — open a user to review or change account options that apply to that member.
- View user details — use the table columns (limits, team, allowed providers, compression status, and related fields) or open a user row for the full record.
If a user has accepted an invitation but is not displayed in the list, click Refresh in the page header to reload the table data.
Teams tab#
Use the Teams tab to create shared groups for limits, compression, and member scoping. Open AI Access → Teams, then:
- Create and manage teams — click + ADD to create a team, Edit to update it, and Delete to remove it.
- Configure team limits — in the Limits column, open the pencil control to set budget, tokens, or requests for the team.
- Enable or disable context compression — use the Enable context compression toggle on the team row.
- Manage team settings — open a team to review or change the shared options that apply to its members.
- View team details — use the table columns (name, limits, compression status, and related fields) or open a team row for the full record.
After teams exist, assign users to a team on the Users tab.
User onboarding workflow#
The following steps illustrate the typical user onboarding workflow in OptScale AI.
To grant a user access to OptScale AI, complete the following steps:
1. Invite the user.
- Go to AI Access → Users.
- Click Invites → Invite.
- Send an invitation to the user.
2. Configure the user's access. After the user accepts the invitation, configure one or more of the following settings:
- Set usage limits.
- Assign the user to a team.
- Configure allowed providers.
- Grant access to the MCP servers.
- Enable or disable context compression, as needed.
3. The user can now sign in and use OptScale AI.
Invite Management page#
The Invite Management page allows administrators to monitor and manage user invitations that have not yet been accepted.
To open the Invite Management page go to AI Access and click INVITES.
The page provides the following information:
- The total number of users with Pending invitations. See summary cards.
- A list of users who have been invited but have not yet accepted their invitation.
- Details of each invitation.
- Actions for managing invitations, such as resending an invitation.
For information about roles and invitation restrictions, see AI Access principles.
Use case: Resend an invitation#
If a user did not receive or has lost their invitation email, you can resend the invitation.
- Go to AI Access → Invites.
- Locate the user in the invitation list.
- In the Actions column, click Resend.
A new invitation email is sent to the user.
Export user data#
Click DOWNLOAD to open a dropdown and choose an export format:
- XLSX spreadsheet — Download the user list as an Excel-compatible file for reporting or offline review.
- JSON file — Download the same data as JSON for automation, backups, or integration with other tools.
The export reflects the users currently visible in the table (respecting any active Search filter).
Invite users#
To invite new users to the organization:
- Open AI Access and click + INVITES to open Invite Management.
- Click + INVITE.
-
On the Invite Users page, enter email and set roles.
- Email: enter one or more emails to invite users to your organization.
- Add role: choose a role.
-
Click Invite.
The new invitation appears in the Invite Management table with Sent and Expires timestamps. After the user accepts, they appear on the main AI Access user list—complete Configure limits, assign the user to a team, and Allowed providers separately.
Configure limits#
Usage limits help administrators control AI consumption by defining maximum budgets, token usage, and request counts for users and teams. Limits can be configured to align with organizational policies, prevent unexpected costs, and ensure fair allocation of AI resources across the organization.
To view or configure limits, go to AI Access, select the Users or Teams tab, and locate the Limits column for the required user or team.
For newly added users, the Limits column is empty until an administrator configures at least one limit.
To configure limits:
-
Open AI Access, select the Users or Teams tab, and locate the Limits column for the required user or team (use Search if needed).
-
In the Limits column, click the
pencil icon. -
In the Edit limits dialog, configure one or more of the following limits:
-
Budget — Specifies the maximum spending limit and its reset interval.
- Tokens — Specifies the maximum number of tokens allowed during each reset period and the reset interval.
-
Requests — Specifies the maximum number of API requests allowed during each reset period and the reset interval.
-
Click SAVE to apply the changes, or CANCEL to discard them.
-
Optional: Click CLEAR LIMITS to remove all configured limits for the selected user or team.
After the limits are saved, the Limits column displays the configured limits, reset intervals, and current usage.
Assign a user to a team#
The Team column lets you assign users to teams for access isolation and resource scoping. For newly added users, the column displays Unassigned (-) until a team is assigned.
-
Open AI Access and locate the user (use Search if needed).
-
In the Team column, click the
pencil icon. -
In the Assign team dialog, select a team from the Team drop-down list, or select Unassigned to remove the user from a team.
-
Click SAVE to apply the assignment, or CANCEL to discard your changes.
After the assignment is saved, the Team column displays the selected team. To change or remove the assignment, click the
pencil icon again and select a different team or Unassigned.
Before assigning users, create and configure teams on the TEAMS tab.
Allowed providers#
Allowed providers control which registered provider connections a user can use. Registering a provider under Providers makes it available to the organization; assigning allowed providers decides which of those connections appear for a specific member in Chat and related access paths (including virtual keys scoped to that user).
When to use it#
- After you add or activate providers, before users open Chat for the first time.
- When users require access to different provider endpoints (for example, one user uses only
ollama, while another uses only managed cloud APIs). - When troubleshooting "provider not in dropdown" reports—confirm that the provider is Active on the Providers page and assigned in Allowed providers for the affected user.
Configure allowed providers#
- Open AI Access and find the user row (use Search if needed).
- In the Allowed providers column, open the edit flow for that user.
- Select one or more Active providers from the organization catalog.
- Save the assignment.
Repeat for each member who needs access. Inviting a user (Invite users) does not assign providers automatically—you configure Allowed providers separately after the account exists in the table, or after a new provider is added.
During initial setup, this step follows Add first provider in First Steps.
Expected behavior#
- Chat — The provider selector shows only providers assigned to the signed-in user (among those that are Active).
- No assignment — If Allowed providers is empty for a user, they will not see unassigned endpoints in Chat even when those providers are healthy for the organization.
Enable context compression#
The Enable context compression setting is available at both the user and team levels.
When context compression is enabled, the platform reduces the size of conversation context sent with each request, which can lower token usage and cost for long or attachment-heavy sessions. When disabled, requests use the full conversation context without compression.
Compression workflow#
The compression engine automatically optimizes request payloads before they are sent to an AI model while preserving the ability to restore the original content when required.
The workflow consists of the following steps:
-
Content routing — The incoming payload is analyzed to determine its content type (for example, JSON, source code, or natural language). The request is then routed to the most appropriate compression algorithm.
-
Type-aware compression — Specialized algorithms optimize each content type using techniques tailored to its structure, reducing token usage while preserving meaning and functionality.
-
Cache alignment — Prompt prefixes are normalized to maximize cache reuse, improving cache hit rates and reducing inference costs and latency.
-
Reversible retrieval — The original content is stored in a local cache and can be restored on demand without affecting request processing.
This workflow reduces token usage, improves cache efficiency, and lowers inference costs while preserving request accuracy.
Configure context compression#
- Open AI Access and select either the Users or Teams tab, depending on whether you want to configure context compression for a user or a team.
- Locate the required user or team (use Search if needed).
- In the Enable context compression column, use the toggle to turn the setting on or off.
The change applies immediately — no separate save action is required.
Expected behavior#
- Enabled — Optimization mode is on for that user or team. Requests may use compressed context and other optimizations when supported for the selected provider and workflow.
- Disabled — Optimization mode is off. Requests use uncompressed context; token volume may be higher for the same conversation length.
- Per user — Each row has its own toggle; team members in the same organization can have different settings.
To review savings from compression and other optimizations, see Optimizations.