RTI Connext Observability Framework
7.7.0
1. Introduction to Connext Observability Framework
1.1. Use Cases
1.2. How Observability Framework Works
1.2.1. Distribution of Telemetry Data
1.2.2. Telemetry Backends
1.2.3. Remote Debugging
1.2.4. Control and Selection of Telemetry Data
1.2.5. Security
1.3. Components
1.3.1. Monitoring Library 2.0
1.3.2. Collector Service
1.3.3. Collector Service Lite
1.3.4. Observability Dashboards
1.4. Telemetry Data
2. Deployments
2.1. Before You Begin
2.2. Evaluation Deployment
2.3. Production Deployments
2.3.1. Single Collector Service instance
2.3.2. Single layer of Collector Service instances
2.3.3. Multiple layers of Collector Service instances
2.3.4. Multiple layers of Collector Service instances with OpenTelemetry Collector
3. Installation
3.1. Monitoring Library 2.0
3.2. Collector Service or Collector Service Lite
4. Usage
4.1. Observability Framework for Production
4.1.1. Monitoring Library 2.0
4.1.2. Collector Service or Collector Service Lite
4.2. Observability Framework for Evaluation
4.2.1. Components used for evaluation
4.2.2. Defining the JSON configuration file
4.2.3. Using the Observability script
4.2.3.1. Create a Docker workspace
4.2.3.2. Initialize and run Docker containers
4.2.3.3. Stop Docker containers
4.2.3.4. Start existing Docker containers
4.2.3.5. Stop and remove Docker containers
4.2.4. Configuring the Docker workspace
4.2.5. Configuring Grafana
4.2.5.1. Initial login
4.2.5.2. Configuration options
4.2.5.3. Create accounts (optional)
4.2.5.4. Change the default time range (optional)
4.2.6. Removing the Observability Framework Docker workspace
5. Monitoring Library 2.0
5.1. Enabling Monitoring Library 2.0
5.2. Setting Initial Metrics and Log Configuration
5.2.1. Enable all metrics
5.2.2. Enable a custom set of metrics
5.3. Configuring Distribution Settings
5.3.1. Setting application name
5.3.2. Changing the default observability domain ID
5.3.3. Setting Collector Service initial peers
5.4. Configuring QoS for Entities
5.5. Connecting to Collector Service Over WAN
6. Collector Service
6.1. Functional Overview
6.1.1. Use cases
6.1.2. Endpoints
6.1.3. Backend integrations
6.1.4. Capability comparison
6.2. Installing a Collector Service
6.3. Using Collector Service
6.3.1. Deploying the executable
6.3.1.1. Running the Collector Service executables
6.3.1.2. Passing command-line arguments
6.3.1.3. Starting the executables using parameters
6.3.1.4. Stopping the executable
6.3.2. Deploying the Docker image
6.3.2.1. Starting the Docker image
6.3.2.2. Stopping the Docker image
6.3.3. Observability Framework evaluation
6.4. Configuration
6.4.1. Builtin configuration profiles
6.4.2. Configuration parameters
6.4.2.1. General parameters
6.4.2.2. WAN parameters
6.4.2.3. Backend integration parameters
6.4.2.4. Security parameters
6.5. REST API Reference
6.5.1. Definitions
6.5.2. Root endpoint (base URL)
6.5.3. API overview
6.5.4. API reference
7. Observability Dashboards
7.1. System Status Dashboards
7.1.1. System Status Dashboard Common Elements
7.1.2. Alert Home Dashboard
7.1.3. Alert Category Dashboards
7.2. Entity List Dashboards
7.3. Entity Status List Dashboards
7.4. Entity Status Dashboards
7.5. Log Dashboards
7.5.1. Log Dashboard
7.5.2. Entity Log Dashboards
7.6. Control Dashboards
7.6.1. Log Control Dashboard
7.6.2. Metric Control Dashboards
7.6.2.1. Single Entity Metric Control Dashboards
7.6.2.2. Multiple Entity Metric Control Dashboards
8. Security
8.1. Secure Communication between Connext Applications and Collector Service or Collector Service Lite
8.2. Secure Communication with Collector Service or Collector Service Lite HTTP Servers
8.2.1. Secure Collector Service HTTP servers (evaluation deployment)
8.2.2. Secure Collector Service or Collector Service Lite HTTP servers (production deployment)
8.2.2.1. Collector Service or Collector Service Lite Docker image
8.2.2.2. Collector Service or Collector Service Lite executable
8.3. Secure Communication with Third-Party Component HTTP Servers
8.3.1. Secure third-party component HTTP servers (evaluation deployment)
8.3.2. Secure third-party component HTTP servers (production deployment)
8.3.2.1. Collector Service Docker image
8.3.2.2. Collector Service executable
8.4. Generating the Observability Framework Security Artifacts
8.4.1. Generating DDS security artifacts
8.4.2. Generating HTTPS security artifacts
8.4.2.1. Preliminary steps
8.4.2.2. Generating a new root CA
8.4.2.3. Generating server certificates
8.4.2.4. BASIC-Auth password file
9. Telemetry Data
9.1. What is Telemetry Data
9.1.1. Levels
9.1.2. Categories
9.2. Resources
9.2.1. Resource Pattern Definitions
9.3. Metrics
9.3.1. Metric Pattern Definitions
9.3.2. Application Metrics
9.3.3. Participant Metrics
9.3.4. Topic Metrics
9.3.5. DataWriter Metrics
9.3.6. DataReader Metrics
9.3.7. Derived Metrics Generated by Prometheus Recording Rules
9.3.7.1. DDS Entity Proxy Metrics
9.3.7.2. Raw Error Metrics
9.3.7.3. Aggregated Error Metrics
9.3.7.4. Enable a Raw Error Metric
9.3.7.5. Custom Error Metrics
9.4. Non-Metric Observables
9.4.1. Application Observables
9.4.2. Participant Observables
9.4.3. Type Observables
9.4.4. Topic Observables
9.4.5. Publisher Observables
9.4.6. DataWriter Observables
9.4.7. Subscriber Observables
9.4.8. DataReader Observables
9.5. Logs
9.5.1. Syslog Levels and Facilities
9.5.2. Activity Context
9.5.3. Log Labels
9.5.4. Collection and Forwarding Verbosity
9.5.4.1. Changing Verbosity Levels Locally
9.5.4.2. Changing Verbosity Levels Remotely
10. Tutorial
10.1. About the Observability Example
10.1.1. Applications
10.1.2. Data Model
10.1.3. DDS Entity Mapping
10.1.4. Command-Line Parameters
10.1.4.1. Publishing Application
10.1.4.2. Subscribing Application
10.2. Before Running the Example
10.2.1. Set Up Environment Variables
10.2.2. Compile the Example
10.2.2.1. Non-Windows Systems
10.2.2.2. Windows Systems
10.2.3. Install Observability Framework
10.2.3.1. Configure Observability Framework for the Appropriate Operation Mode
10.2.4. Start the Collection, Storage, and Visualization Docker Containers
10.3. Running the Example
10.3.1. Start the Applications
10.3.2. Changing the Time Range in Dashboards
10.3.3. Simulate Sensor Failure
10.3.4. Simulate Slow Sensor Data Consumption
10.3.5. Simulate Time Synchronization Failures
10.3.6. Change the Application Logging Verbosity
10.3.7. Change the Metric Configuration
10.3.7.1. Resources used in this example
10.3.7.2. Changing metrics collected for a single DataWriter
10.3.7.3. Changing metrics collected for all DataWriters of an application
10.3.8. Remote Debugging with Admin Console
10.3.9. Close the Applications
11. Troubleshooting Observability Framework
11.1. No Observable Data available on WebSocket or Observable backends
11.1.1. Connext application discovery logs
11.1.2. Collector Service and Collector Service Lite Discovery Logs
11.2. Docker Container[s] Failed to Start
11.2.1. Check for port conflicts
11.2.2. Check that you have the correct file permissions
11.3. No Data in Dashboards
11.3.1. Check that Collector Service has discovered your applications
11.3.2. Check that Prometheus can access Collector Service
11.3.3. Check that Grafana can access Prometheus
11.3.4. Check that Grafana can access Loki
11.4. Can Collector Service run in Windows or macOS?
12. Glossary
13. Release Notes
13.1. Supported Platforms
13.2. Compatibility
13.3. Supported Docker Compose Environments
13.4. Supported Docker Environments for Collector Service
13.5. What’s New in 8.0.0
13.5.1. Support for new observable DataReaderQos Lifespan QoS Policy in Collector Service
13.5.2. New standalone Collector Service Lite executable simplifies observability deployments
13.5.3. New Collector Service location tag identifies the source of monitor data in large deployments
13.5.4. Reduced CPU and memory utilization for Collector Service or Collector Service Lite when using builtin “Forwarder” profiles
13.5.5. Improved remote debugging performance with Admin Console under certain DDS loads
13.5.6. Faster application detection in Admin Console during remote debugging increases overall operational efficiency
13.5.7. Third-Party Software Changes
13.6. What’s Fixed in 8.0.0
13.6.1. Usability
13.6.1.1.
[Critical]
When remote debugging, Admin Console failed to receive data from Collector Service over WebSocket when an application crashed
13.6.1.2.
[Major]
Delayed WebSocket updates when using remote debugging
13.6.1.3.
[Trivial]
WARNING log messages occurred for conditions that are expected under normal operation
13.6.2. Performance and Scalability
13.6.2.1.
[Critical]
Monitoring commands may have experienced long delays and been processed in bursts
13.6.2.2.
[Major]
Forwarded commands may have automatically been purged before Collector Service could process them
13.6.2.3.
[Major]
Monitoring Library 2.0 collected periodic metrics when Collector Service was not running
13.6.3. Logging
13.6.3.1.
[Trivial]
Potential unexpected log messages while disabling Monitoring
13.6.4. Crashes
13.6.4.1.
[Critical]
Collector Service crashed or experienced data corruption during service shutdown
13.6.4.2.
[Critical]
Potential application crash when enabling Monitoring Library 2.0
13.6.4.3.
[Critical]
Potential application crash due to memory reordering when Monitoring Library 2.0 was enabled
13.6.5. Hangs
13.6.5.1.
[Critical]
Requester or Replier creation may have deadlocked when Monitoring Library 2.0 was enabled
13.7. Known Issues
13.7.1. Connext applications may crash when using multiple language bindings
Copyrights and Notices
RTI Connext Observability Framework
Index
Index