nx1-deployer image
v1.17.0Before you upgrade
Check these three changes before you deployv1.17.0.
Mirror the new Spark image names
Every Spark-family image now carries its Spark major version in its name. Mirror the images you need to your registry before you deploy:- For Spark 3.5:
spark3,airflow-spark3,jupyterhub-spark3, andkyuubi-spark3. - For Spark 4.2:
spark4,airflow-spark4,jupyterhub-spark4, andkyuubi-spark4.
Plan your Ranger policy migration
If you changeranger_repo_mode on an existing tenant, then NexusOne creates a new repository that contains only
default policies. Your existing policies don’t carry over. Plan how you’ll migrate them before you switch modes.
See Ranger repository modes.
Check S3 delete permissions
Deleting or moving an object in the JupyterHub S3 browser, or deleting one withs3Cli rm, now requires Ranger delete permission instead of write. A user who could
delete objects with only write permission before can’t after you upgrade.
See JupyterHub S3 now checks for delete permission.
New features
This section contains new features recently added to the NexusOne platform.Spark 4.2
A Spark 4.2 image is now available alongside the existing Spark 3.5 image. Choose between them with the newspark_version tenant variable.
Pass "3" for Spark 3.5.6 on Scala 2.12, or "4" for Spark 4.2.0 on Scala 2.13. Defaults to "3".
Airflow, JupyterHub, and Kyuubi each match the version you choose automatically, so the whole tenant stays on one
consistent Scala version.
Spark Connect through Kyuubi
Kyuubi can now run an optional Spark Connect frontend on port15002. It lets Spark Connect clients reach a tenant’s
Spark without a full Spark install on the client side.
Set the tenant variable spark_connect_enabled to true to enable it. Defaults to false. Spark Connect only works on
Spark 4, so set spark_version to "4" too.
Spark configuration and catalog discovery
Spark used to get its configuration fromcopysparkhome and the utils mount. This release replaces both with a new
spark-conf module, which delivers Spark’s configuration through a Kubernetes Secret and its own dedicated Keycloak
client.
Spark also now discovers Iceberg REST catalogs, including Gravitino, at runtime through AiAPI. Before, each catalog
needed a static entry in spark-defaults.conf.
NX1 Decisions
NexusOne now discovers decision workspaces automatically and lets you query them through SQL functions in both Trino and Spark. The tenant variabledecisions_enabled controls this feature. Defaults to true.
Distributed tracing for Spark and Kyuubi
You can now send OpenTelemetry traces from Spark and Kyuubi to Tempo. This lets you follow a single query from Kyuubi through to the Spark jobs it runs. Set the tenant variabletracing_enabled to true to enable it. Defaults to false. Tracing requires the tenant’s
embedded monitoring stack.
Centralized Grafana for the modules tier
Each tenant already has its own Grafana. This release adds an optional, platform-level Grafana in the modules tier. It gives you one view across every tenant’s metrics, logs, and traces, instead of a separate login per tenant. Central Grafana queries each observability-enabled tenant’s existing Mimir, Loki, and Tempo backends directly. It doesn’t change where any tenant’s telemetry is stored. Enable it in your modules tier config by using the following new variables:central_grafana_enabled:trueto deploy the central Grafana. Defaults tofalse.central_grafana_sync_enabled:trueto deploy the datasource sync that discovers tenants. Only used whencentral_grafana_enabledistrue. Defaults totrue.
- A per-namespace workload table: Pods running, pending, and failed, plus restarts and containers not ready.
- CPU and memory requested vs used.
- A table of Kubernetes events, with a count of warning events in the last hour.
Ranger repository modes
A new tenant variable,ranger_repo_mode, sets how NexusOne groups a tenant’s Ranger policies into repositories. It
accepts three modes:
tenant: The tenant gets its own repository.shared: The tenant shares one repository with other tenants.group: The tenant shares a repository with a group of tenants.
single_ranger_repo variable still works as an alias.
This release also adds support for running mixed Ranger releases, optional cleanup of URL policies, and cleanup of
unused service definitions.
Dedicated Hive Metastore database
This release adds a new tenant variable,metastore_dbname. Set it to point a tenant’s Hive Metastore to a
dedicated Postgres database, instead of the shared hive schema in the core database.
Create this database on the Postgres instance set in core tier’s dbhost variable. A tenant inherits that connection
automatically, so you don’t set it yourself.
If you leave metastore_dbname empty, then NexusOne keeps using the shared hive schema.
Karpenter placement for Trino and Spark
You can now let Karpenter place Trino workers and Spark executors, so a cluster scales its nodes to match query and job load. Set the tenant variablekarpenter_enabled to true to enable it. Defaults to false.
Superset MCP service
Superset now includes an MCP service, integrated with AiAPI, so NexusOne’s AI features can work with Superset directly.Airflow DAG bundles and Keycloak sign-in
Airflow now supports dedicated DAG bundles, and you can sign in to Airflow directly through Keycloak.Use Claude through your AWS account
NexusOne’s AI features, such as crews and natural-language SQL, could already call Claude through the public Anthropic API. This release lets you reach Claude through Anthropic’s own Claude Platform hosted on AWS instead. This is important if your AI traffic needs to stay inside your own cloud for compliance reasons. Configure it in NexusOne’s Tenant Manager with these new tenant variables:llm_anthropic_base_url: Overrides the base URL for the Anthropic API, so you can reach Claude through your Claude on AWS API, instead of Anthropic’s standard one. Leave empty to use Anthropic’s standard API endpoint.llm_anthropic_workspace_id: Workspace ID for Claude on AWS.llm_extra_headers: If your LLM gateway or backend needs custom authentication headers, then set them here. AiAPI attaches them to every LLM call.
- The NX1 LLM router now accepts portal JWTs for authentication.
- A new Trino
aicatalog connector.
New Keycloak roles
The following roles are now available in NexusOne and you can assign them to users:- Centralized Grafana roles:
nx1_grafana_platform_admin: Sign in to the central Grafana with the Admin rolenx1_grafana_platform_viewer: Sign in to the central Grafana with the Viewer role
- Engineering roles:
nx1_engineer_admin: Engineering administrator access
- JupyterHub roles:
nx1_jupyterhub_elyra: Turns on Elyra in JupyterHub for a tenant’s users
Bug fixes
This section contains fixes for issues affecting apps or features on the NexusOne platform.Spark scratch and upload paths are now scoped per tenant
When tenants shared a Ranger repository, Spark’s S3 scratch and upload paths weren’t scoped to each tenant. That meant one tenant could reach another tenant’s scratch and upload data. This release scopes both paths per tenant, so each tenant can only reach its own. This release also fixes the Spark history UI’s session-key handling and its OpenLineage URL.JupyterHub S3 now checks for delete permission
Deleting or moving an object in the JupyterHub S3 browser, or deleting one withs3Cli rm, only checked for Ranger write permission. A user who could write to a bucket
could also delete from it.
This release requires delete permission for those actions instead. See
Check S3 delete permissions before you upgrade.
Kyuubi stops refetching Keycloak’s signing keys on every request
Kyuubi used to rebuild its whole authentication cache from scratch on every single request, not just after a restart. This meant every request refetched Keycloak’s signing keys, and an outage-tolerance grace cache built for exactly this case never actually held anything. This release shares that cache across requests, so Kyuubi reuses a signing key it already fetched, instead of refetching it every time. This release also corrects how Kyuubi exposes the Spark Connect port.Deployment reliability
A tenant apply could time out waiting on a PVC that never left thePending state. In this release, NexusOne no
longer waits for tenant PVCs to bind before continuing.
Other fixes
- DataHub: Fixed a startup crash. Every DataHub component now receives
DATAHUB_SYSTEM_CLIENT_SECRET. - Airflow: Fixed Keycloak signing-key resolution, so Airflow now retrieves the key on each login.
- AiAPI: Fixed authentication for tokens generated through
admin-cli. - Envoy: Fixed 404 responses for router-rewritten paths outside
/v1/. - Grafana: Fixed Terraform plans that kept showing a diff for alert rule groups on every run.
- Ranger sync: Two versions of the sync agent no longer run at the same time during an upgrade.
- Superset: Fixed initialization on the hardened image, which doesn’t include coreutils.
- JupyterHub: Fixed the notebook pod’s
fsGidon OpenShift when you use shared, static home storage.
Enhancements
This section contains enhancements to existing app features on the NexusOne platform.Platform component upgrades
This release upgrades these platform components:- Trino to 483
- Kyuubi to 1.12
- Grafana Mimir to 3.2
- Keycloak to 26.7
- Airflow from 3.2.2 to 3.3, with the image tag
3.3.11 - DataHub to v1.7.0.1
Airflow parallelism
Airflow’s parallelism and its maximum active runs per DAG both increase to128.
DataHub
Besides the upgrade to v1.7.0.1, this release changes DataHub in these ways:datahub-upgradereplaces the legacy setup jobs.- Telemetry is turned off.
- Kafka topics, new and existing, now use zstd compression and a 50 MB message limit.
- The business attributes and metrics features are turned on.
create_networkpolicy to true to create them.
Defaults to false.
JupyterHub
These JupyterHub settings are now configurable per tenant:- Session timeout: Set the tenant tier variable
jupyter_session_timeoutto how long, in seconds, an idle notebook server and its kernel run before NexusOne shuts them down. Defaults to3600. - Notebook image tags: Choose which notebook image tags a tenant uses.
- Elyra: Turn on Elyra for a tenant’s users with the new
nx1_jupyterhub_elyraKeycloak role.
Grafana content controls
nx1-deployer imagev1.16.1 added the tenant tier variable grafana_alerts_enabled. It only turned off NexusOne’s
built-in Grafana alert rules for platform components like Spark and Gravitino. This release adds a dedicated Grafana
content module, controlled by a broader variable:
grafana_content_enabled: Whether NexusOne manages Grafana’s content or not. If you disable it, then you manage that content yourself, while Mimir, Loki, and Tempo keep running unaffected. Defaults totrue.
grafana_alerts_enabled variable is now a subset of this new variable, so it’s ignored
when grafana_content_enabled is false.
This release also adds Grafana dashboard embedding in the portal, and support for collection tiers that span
namespaces.
Kyuubi NX1 PSK authentication
Kyuubi can now authenticate you with a personal NX1 user PSK or Keycloak JWTs. PSK authentication is on by default. For a personal NX1 user PSK, Kyuubi caches resolved identities, so it doesn’t call back to the NX1 API on every request. You can also now sign in to the Kyuubi Web UI through Keycloak.S3 endpoint formats
You can now specify an S3 endpoint as eitherhost or scheme://host. If you don’t set a region, then it defaults to
the cluster’s region.
Secrets in deployment metadata
Every Terraform variable that holds a secret is now marked sensitive, and the deployer includes the patched Helm provider3.3.0. Together, these keep secrets out of deployment metadata. This release also removes the unused Vault
provider.
Scheduling and ingress
- YuniKorn: When you enable YuniKorn, it now schedules the remaining pods too, not just some of them.
- Envoy: New configuration options for the Envoy listener section, and for how Envoy handles the HTTPS forwarded-proto header.
- Trino: Trino no longer logs the AWS SDK v1 deprecation notice.
Security
This release updates the following components with hardened or CVE-patched images:- Alloy
- JupyterHub configurable-http-proxy
- Keycloak
- Loki
- oauth2-proxy
- OpenSearch
- postgres-exporter
- Ranger and ranger-sync
- Redis
- Superset
- The observability router and webhook
- YuniKorn scheduler
Image versions
The following images changed in this release. Images not listed keep theirv1.16.1 versions.

