Release Notes
Current version: 1.5.12.
Upgrade method:
-
Docker Compose: update the image tag in
x-ops-imageat the top ofops.yaml, then run:docker compose -f ops.yaml pulldocker compose -f ops.yaml up -d -
Kubernetes: download the deployment manifests for the target version again and apply them over the existing resources.
Upgrade does not modify existing data volumes. Monitoring history and alert configuration are retained by default.
Read before upgrading from 1.4.x
- Data source configuration: data sources are now maintained on the UI "Data Sources" page.
ENV_MYSQL_*,ENV_REDIS_*,ENV_KAFKA_ENDPOINTS,ENV_ELASTICSEARCH_*,ENV_MONGODB_URI,ENV_PROMETHEUS_HOST,ENV_FLINK_URL, and similar variables inops.yaml/ConfigMap are no longer read. Existing data sources are not affected, and these historical variables can be removed. - Object storage: logs and traces must use two different buckets,
ENV_S3_BUCKET_LOKIandENV_S3_BUCKET_TEMPO, and the buckets must be created before deployment. - Data directory: the
/data/mdis/layout applies only to new installations. Do not directly change existing data volume paths during upgrade.
v1.5 · July 2026
This series consolidates collection configuration into a single entry point and improves Kubernetes monitoring, multi-cluster collection, and object storage capabilities.
Added
- YAML/JSON batch import and export for data sources, supporting batch onboarding, backup/restore, and cross-environment migration.
- Connection tests now also check collection status, distinguishing reachable connections, active collection, and reachable connections with no data collected yet.
- Kubernetes collection-component mode. Monitored clusters can push metrics back to the Ops Platform through
remote_write, suitable for scenarios where the Ops Platform is outside the monitored cluster or manages multiple clusters. - Gateway endpoints with Bearer authentication for metric and log writes, used for multi-cluster metric reporting and cross-cluster log push.
- Custom pages: register any Web UI (such as Flink or Langfuse) into the "Custom" sidebar group, with two embedding modes (reverse proxy or direct iframe). The previous Flink entry is migrated automatically on upgrade; the menu and URL stay unchanged.
- Multi-node MySQL/Redis/Elasticsearch data sources: each node is collected separately, so Redis Sentinel/Cluster and multi-node ES deployments can be registered as one source.
- Kafka lag alerts can target an entire consumer group.
- Log panels load more entries automatically when scrolled to the bottom.
Changed
- Data source configuration is unified on the UI "Data Sources" page. Historical data source variables in
ops.yaml/ConfigMap no longer take effect. - Kubernetes monitoring now uses in-cluster collection: deploy the Ops Platform inside the monitored cluster, or install collection components in the monitored cluster.
- Object storage for logs and traces is split into independent buckets. Component data directories are unified as
/data/mdis/<component>. - Namespace, node label, and taint key are unified as
hap-ops. - The "Flink" data source type is merged into "Custom pages"; configuration export/import now covers custom pages.
- The Kafka rebalance alert no longer offers an automatic restart option.
Key Fixes
- Fixed empty monitoring data after Kubernetes deployment, missing filter options on Kubernetes panels, and missing host monitoring data.
- Fixed metric collection issues caused by MongoDB multi-node replica sets, credential passing, and insufficient permissions.
- Fixed slow-query diagnostics hot reload, permission hints, and inaccurate exception classification.
- Fixed inaccurate collection status display, invisible Flink connectivity state, and missing prefix in
remote_writeexample addresses. - Fixed Kubernetes panels showing no data at the node level (cadvisor metrics lacked the
nodelabel) and collection components being evicted under Kubernetes due to missing ephemeral-storage requests. - Fixed the alert engine not being initialized, which made rules ineffective, "Check now" fail, and alert history stay empty.
- Fixed Docker Compose deployments in a standalone directory not joining the product network, so middleware could not be reached by service name.
- Fixed the service-log self-check reporting "not configured" when logs were pushed but rejected by Loki; it now reports the discard reason (such as clock skew).
Patch version summary (1.5.0 - 1.5.12)
| Version | Highlights |
|---|---|
| 1.5.12 | Custom pages (Flink entry generalized) Export/import covers custom pages |
| 1.5.11 | Alert engine initialization fix Per-node collection for multi-node sources Kafka consumer-group lag alerts Log panel infinite scrolling Compose network fix |
| 1.5.10 | Kubernetes panel node-level fix Ephemeral-storage requests for collection components |
| 1.5.9 | Host panel fixes Data source export Connection test collection-status validation |
| 1.5.8 | Kubernetes cluster identifier injection Slow-query target hot reload Flink connectivity status |
| 1.5.7 | Kubernetes single-file manifest adds RBAC/kube-state-metrics Bearer write endpoints |
| 1.5.6 | Multi-node replica set MongoDB collection |
| 1.5.5 | Kubernetes collection-component mode MongoDB collection credential passing fix |
| 1.5.4 | On-demand Docker log collection under Kubernetes Cluster manifest download |
| 1.5.3 | Kubernetes monitoring consolidated to in-cluster collection |
| 1.5.2 | Data directory unified as /data/mdis/<component> |
| 1.5.1 | Separate object storage buckets for logs and traces |
| 1.5.0 | Unified data source entry point and batch import |
v1.4 · July 2026
This series introduces the alert subsystem and changes delivery to a single-image form.
Added
- Alert subsystem supporting rule types for basic resources, ports, SSL, MySQL, MongoDB, Kafka lag, Kafka rebalance, and more.
- "Data Sources" page for centralized management of host, middleware, Kubernetes, and Flink data sources.
- Collection status and deployment self-check capabilities, helping locate collection issues from the UI.
- MongoDB slow-query analysis now supports sharded clusters and multi-instance aggregation.
Changed
- All components are merged into the
ops-allinoneimage, with runtime roles distinguished byROLE. - Alert rules, status, history, notification channels, and data source configuration are stored independently in
ops-mongo. - Added required credential encryption key
ENV_ALERT_CRYPTO_KEY. - Historical Grafana alert, notification, and silence configuration is replaced by the Ops Platform alert capability.
Key Fixes
- Fixed host monitoring loss after upgrade, incomplete default values when creating data sources through API, and duplicate data sources.
- Fixed MySQL/MongoDB credential decryption, object storage compatibility, and
ops-mongostartup order dependency issues. - Fixed tracing service-name matching that previously worked only for the
defaultnamespace.
Patch version summary (1.4.0 - 1.4.10)
| Version | Highlights |
|---|---|
| 1.4.10 | Host monitoring upgrade-loss fix API data source defaults Deployment self-check |
| 1.4.8 | Data source collection status |
| 1.4.6 | Tracing service names support any namespace |
| 1.4.5 | ops-mongo background retryData sources enabled by default after creation |
| 1.4.4 | Prometheus retention by size |
| 1.4.3 | Object storage compatibility Istio tracing appProtocol |
| 1.4.1 | MySQL/MongoDB credential decryption |
| 1.4.0 | Alert subsystem, unified data sources, and single-image delivery |
v1.3 · June 2026
- Added Kubernetes cluster monitoring dashboards and menu.
- Supported collecting multiple middleware instances with a single agent.
- Loki supports retention policies and delete API.
v1.2 · April-May 2026
- Added log search and APM distributed tracing capabilities.
- Split log panels into container console and service logs, with structured field expansion, full-text search, and multi-condition filtering.
- Supported sub-path deployment.
- Deployment resource requirements changed to 4C/8G.
v1.1 · February 2025
- Added MongoDB slow-query and index suggestion capabilities.
v1.0 · January 2025
- Initial version with system resource monitoring.