Menu

AI-Controlled Cloud Phone API

DuoPlus Cloud Phone Automation Control API

This API provides a service exposed from within the cloud phone; it allows for direct control of the cloud phone and the execution of automated operations, such as: Get the Screen and UI Elements、Tap and Long Press、Text Input and Keyboard、Swipe and Scroll、App and Page Navigation、Wait and Read Elements、etc.

Prerequisites: Get an API Key from Automation > API in the DuoPlus console. The target phone must be running (status == 1) with HTTP enabled (http_status is 1 or "1"). If it is stopped, start it in the console or via the platform power-on API; this document does not cover that endpoint. The Gateway uses the same API Key as its Bearer Token. Device-side authentication must be provisioned to match it, or the Gateway may return 401.

1. Get the Cloud Phone IP, Region, and Automation Support Status

Before calling the HTTP Gateway, use the platform list endpoint to obtain the target cloud phone's ip and region, and confirm that status == 1 and http_status is 1.

1.1 Endpoint

text Copy
POST https://openapi.duoplus.net/api/v1/cloudPhone/list

Request headers:

http Copy
Content-Type: application/json
Lang: zh
DuoPlus-API-Key: <API_KEY>

1.2 Request Parameters

Query by exact cloud phone ID:

json Copy
{
  "image_id": ["IMAGE_ID"],
  "page": 1,
  "pagesize": 100
}

You can also query by name:

json Copy
{
  "name": "phone-name",
  "page": 1,
  "pagesize": 100
}
Parameter Required Description
page No Page number, starting at 1; defaults to page 1
pagesize No Number of items per page; maximum 100, default 10
image_id No Array of cloud phone IDs; an exact ID is recommended
name No Cloud phone name; names may not be unique
link_status No Array of connection statuses
group_id No Group ID
proxy_id No Proxy ID

1.3 Response Fields Required for Automation

Example response envelope:

json Copy
{
  "code": 200,
  "message": "success",
  "data": {
    "list": [
      {
        "id": "IMAGE_ID",
        "name": "phone-name",
        "status": 1,
        "ip": "10.0.0.10",
        "region": "sg",
        "http_status": 1
      }
    ],
    "total_page": 1
  }
}

This example shows only the fields needed for automation. The server may return additional device information.

Field Purpose
data.list[].id Cloud phone ID, used to confirm that the correct device was selected
data.list[].status 1 means the cloud phone is running; handle other states first
data.list[].ip Value of the Gateway CloudIP request header
data.list[].region Value of the Gateway Region request header; use the selected phone's returned value
data.list[].http_status 1/"1" means HTTP is enabled; 0/"0" means it is disabled
data.total_page Total number of pages; iterate through all pages when not querying by exact ID

Success requires both a successful HTTP request and a top-level response code == 200.

1.4 cURL Example

bash Copy
curl -sS 'https://openapi.duoplus.net/api/v1/cloudPhone/list' \
  -H 'Content-Type: application/json' \
  -H 'Lang: zh' \
  -H "DuoPlus-API-Key: $DUOPLUS_API_KEY" \
  --data '{"image_id":["IMAGE_ID"],"page":1,"pagesize":100}'

Verify the matching item's id, then extract:

text Copy
CloudIP = matched_phone.ip
Region  = matched_phone.region

Call the automation Gateway only when status == 1 and String(http_status) == "1". A nonempty ip does not mean automation is enabled. When searching by name, read every page indicated by total_page; use the ID to resolve duplicate names. The platform OpenAPI limit is 1 QPS per endpoint, including pagination requests.

2. Automation Endpoint

Send all automation operations to:

text Copy
POST https://agent-gateway.duoplus.net/agent-command
http Copy
Content-Type: application/json
Authorization: Bearer <API_KEY>
Region: <REGION>
CloudIP: <CLOUD_IP>
Header Description
Authorization Gateway authentication, in the form Bearer <API_KEY>
Region The target phone's region returned by list
CloudIP Cloud phone private IP address

The cloud phone must be running with HTTP enabled. If the Gateway returns 401, check the API Key and device authentication configuration rather than retrying blindly. Replace the example region and IP in the Section 13 cURL commands with the values returned for the target phone.

3. Common Request Model

operation Purpose
health Check Gateway/backend health
ready Check whether the automation executor is ready
submit Get the screen/UI or perform a UI action
query Query a retained command result
stop Stop a running or stuck task

General request for an action:

json Copy
{
  "operation": "submit",
  "command_id": "click-element-1730000000000-a1b2c3d4",
  "task_id": "click-element-task-1730000000000-a1b2c3d4",
  "action": "execute",
  "payload": {
    "task_type": "ai",
    "task_id": "click-element-task-1730000000000-a1b2c3d4",
    "action": "execute",
    "action_name": "CLICK_ELEMENT",
    "params": {
      "text": "Search",
      "wait_after": 500
    }
  }
}
  • command_id and task_id should be unique.
  • Use uppercase action names for action_name.
  • Put action parameters in payload.params.
  • You may add a Unix timestamp in milliseconds as a top-level deadline_at.
  • A synchronous action can take up to approximately 300 seconds; set the client timeout to approximately 310 seconds.

The params for any action may include:

Parameter Description
wait_before Delay before the action, in milliseconds
wait_after Delay after the action, in milliseconds; 500-1500 is recommended for navigation/loading

4. Health and Readiness Checks

Health check:

json Copy
{"operation":"health"}

Readiness check:

json Copy
{"operation":"ready"}

The executor is ready when the response has executor == "ready" or ready == true. On 503 or NOT_READY, continue polling within the startup timeout.

5. Get the Screen and UI Elements

5.1 Request

json Copy
{
  "operation": "submit",
  "command_id": "ui-state-1730000000000-a1b2c3d4",
  "task_id": "ui-state-task-1730000000000-a1b2c3d4",
  "action": "get_ui_state",
  "payload": {
    "task_type": "ai",
    "task_id": "ui-state-task-1730000000000-a1b2c3d4",
    "action": "get_ui_state",
    "lang": "zh"
  }
}

5.2 Parse the Response

result_json is a JSON string and must be parsed separately:

javascript Copy
const outer = await response.json();
const result = JSON.parse(outer.result_json);

The parsed object may contain:

  • success: whether the action succeeded;
  • a UI element tree, used to locate controls by text, resource ID, content description, and other attributes;
  • screenshot: a Base64-encoded image of the current screen.

Decode the screenshot:

javascript Copy
const encoded = result.screenshot.replace(/^data:[^,]+,/, "");
const bytes = Buffer.from(encoded, "base64");

Read the UI before the first action. Read it again after every click, input, swipe, or navigation action to verify the actual screen change.

6. Tap and Long Press

6.1 Tap an Element: CLICK_ELEMENT

json Copy
{"text":"Continue","wait_after":800}
Parameter Description
resource_id Android resource ID; usually the most stable selector
content_desc Accessibility content description
text Displayed element text
class_name Android view class name
element_order Zero-based index when multiple elements match
json Copy
{"resource_id":"com.example:id/login"}
json Copy
{"content_desc":"Search"}
json Copy
{"text":"Item","element_order":1}

Prefer resource_id, then content_desc/text, then coordinates.

6.2 Long Press an Element: LONG_ELEMENT

json Copy
{"text":"Message","duration":1200}

Use the same selector fields as CLICK_ELEMENT. duration is in milliseconds.

6.3 Coordinate Actions

Tap with CLICK_COORDINATE:

json Copy
{"x":500,"y":420,"wait_after":500}

Long press with LONG_COORDINATE:

json Copy
{"x":500,"y":420,"duration":1200}

Double tap with DOUBLE_TAP_COORDINATE:

json Copy
{"x":500,"y":420}

Coordinates are relative values in the range 0..1000, with (0,0) at the top left and (1000,1000) at the bottom right:

text Copy
relative_x = round(pixel_x / screenshot_width  * 1000)
relative_y = round(pixel_y / screenshot_height * 1000)

Take a fresh screenshot immediately before a coordinate action and use the target's center. Do not reuse old coordinates after a page transition or rotation.

7. Text Input and Keyboard

7.1 Enter Text: INPUT_CONTENT

json Copy
{"content":"hello world","clear_first":true}

Focus the input field first: get the UI, tap the input field, call INPUT_CONTENT, send Enter if needed, and get the UI again to verify the result.

7.2 Keyboard: KEYBOARD_OPERATION

json Copy
{"key":"enter"}

Supported keys: enter, delete, tab, escape, and space.

8. Swipe and Scroll

Action name: SLIDE_PAGE.

Default gesture:

json Copy
{"direction":"up","wait_after":600}

Supported directions: up, down, left, and right. To scroll down through content, swipe your finger upward and use up.

Custom gesture path:

json Copy
{
  "direction": "up",
  "start_x": 520,
  "start_y": 780,
  "end_x": 500,
  "end_y": 260,
  "wait_after": 600
}

Try the default gesture first and reread the screen after every swipe. If it has no effect, adjust the path slightly; generally try the same target no more than three times.

Open an app with OPEN_APP:

json Copy
{"package_name":"com.android.settings","wait_after":1200}

Go to the home screen with GO_TO_HOME:

json Copy
{}

Go back with PAGE_BACK:

json Copy
{}

10. Wait and Read Elements

Wait a fixed duration with WAIT_TIME:

json Copy
{"wait_time":1500}

Wait for an element with WAIT_FOR_SELECTOR:

json Copy
{"resource_id":"com.example:id/result","timeout":15}

Or:

json Copy
{"text":"Completed","timeout":10}

timeout defaults to 10 seconds.

Read the text of one element with GET_SINGLE_ELEMENT_TEXT:

json Copy
{"resource_id":"com.example:id/title"}

You can use selector fields such as text, resource_id, and class_name.

11. Query Results and Stop Tasks

Query a command result:

json Copy
{"operation":"query","command_id":"<COMMAND_ID>"}

Stop a stuck task:

json Copy
{
  "operation": "stop",
  "command_id": "stop-1730000000000-a1b2c3d4",
  "task_id": "<TASK_ID>",
  "reason": "operation timeout"
}

Use stop only for long-running or stuck tasks, not for operations that have completed normally.

12. Determine Success

For a submit UI action, HTTP 200 alone does not indicate success. All of the following must hold:

  1. The HTTP request succeeds.
  2. The outer response has state == "SUCCEEDED".
  3. After parsing, result_json.success == true.
  4. A fresh UI read shows the expected screen state.
javascript Copy
const outer = await response.json();
if (outer.state !== "SUCCEEDED") throw new Error("Gateway command failed");

const result = JSON.parse(outer.result_json);
if (result.success !== true) throw new Error("UI action failed");

// Call get_ui_state again to verify the actual UI result.

13. Complete cURL Examples

13.1 Get the Current Screen

bash Copy
curl -sS 'https://agent-gateway.duoplus.net/agent-command' \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $DUOPLUS_API_KEY" \
  -H "Region: $DUOPLUS_REGION" \
  -H "CloudIP: $DUOPLUS_CLOUD_IP" \
  --data '{
    "operation":"submit",
    "command_id":"ui-state-UNIQUE_ID",
    "task_id":"ui-state-task-UNIQUE_ID",
    "action":"get_ui_state",
    "payload":{
      "task_type":"ai",
      "task_id":"ui-state-task-UNIQUE_ID",
      "action":"get_ui_state",
      "lang":"zh"
    }
  }'

13.2 Tap an Element

bash Copy
curl -sS 'https://agent-gateway.duoplus.net/agent-command' \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $DUOPLUS_API_KEY" \
  -H "Region: $DUOPLUS_REGION" \
  -H "CloudIP: $DUOPLUS_CLOUD_IP" \
  --data '{
    "operation":"submit",
    "command_id":"click-UNIQUE_ID",
    "task_id":"click-task-UNIQUE_ID",
    "action":"execute",
    "payload":{
      "task_type":"ai",
      "task_id":"click-task-UNIQUE_ID",
      "action":"execute",
      "action_name":"CLICK_ELEMENT",
      "params":{"text":"Continue","wait_after":800}
    }
  }'

Before running the Section 13 examples, set DUOPLUS_API_KEY, DUOPLUS_REGION, and DUOPLUS_CLOUD_IP. Replace every occurrence of UNIQUE_ID with a fresh unique value for each request; the top-level and payload task_id values must match within one request. These commands use Bash syntax; PowerShell environment variable syntax differs.

14. Automation Recommendations

  • Read the UI before each action, then read it again and verify the result afterward.
  • Prefer element selectors. Use coordinates for WebViews, Canvas content, or controls that cannot be identified.
  • Add approximately 500-1500 ms of wait_after for navigation or loading actions.
  • Retry the same approach no more than three times. If nothing changes, switch selectors, coordinates, or navigation methods.
  • Pause for user input before CAPTCHAs, payments, account recovery, or destructive confirmations.
  • A screenshot shows what happened; it does not by itself prove that the task is complete.

15. Source

Previous
How to permanently enable the accessibility permission for the app.
Next
Cloud Phone RPA Documentation
Last modified: 2026-09-28Powered by