Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
62 commits
Select commit Hold shift + click to select a range
93bef6b
Updated NodeNorm loader to v2.3.27 and Babel 2025nov1.
gaurav Nov 6, 2025
f5c9701
Upgraded NodeNorm Web to v2.3.27.
gaurav Oct 30, 2025
ae2af70
Switched translator-exp to load.
gaurav Nov 10, 2025
3f4f863
Updated NameRes Solr to v9.
gaurav Nov 10, 2025
1a55b11
Updated NameRes to 2025nov4.
gaurav Nov 10, 2025
13d426f
Upgraded NodeNorm to 2025nov4.
gaurav Nov 10, 2025
29edb39
Updated NodeNorm renci-exp to v2.3.27.
gaurav Nov 11, 2025
b494a45
Upgraded backup scripts to 2025nov4.
gaurav Nov 14, 2025
a09d608
Reverted compressed file transfers.
gaurav Nov 14, 2025
d76900a
Updated compressed download for db3.
gaurav Nov 14, 2025
fc72834
Added NodeNorm Exp to whitelist.
gaurav Nov 16, 2025
c7270f9
Merge branch 'update-nodenorm-nameres-2025nov4' of github.com:helxpla…
gaurav Nov 16, 2025
d8d7c03
Added taxon_specific field.
gaurav Nov 17, 2025
20ed899
Upgraded NodeNorm to v2.3.28.
gaurav Dec 4, 2025
91437c4
Updated NodeNorm Loader to v2.3.28.
gaurav Dec 4, 2025
03c97dd
Upgraded NameRes to v1.6.0.
gaurav Dec 4, 2025
8549953
Updated NodeNorm Loader to v0.8.0.
gaurav Dec 8, 2025
c675a15
Switched NodeNorm Loader Dev to restore.
gaurav Dec 8, 2025
47b0f32
Replaced redis-rdb-cli with new reader.
gaurav Dec 8, 2025
912df0e
Replaced tag with the development code.
gaurav Dec 8, 2025
fec5cc0
Attempted to add redis_version everywhere.
gaurav Dec 8, 2025
94d5a5a
Changed NodeNorm Exp to use the Dev database.
gaurav Dec 8, 2025
7d0458f
Corrected redis_config.redis_version.
gaurav Dec 8, 2025
a0a280d
Switched NodeNorm Exp to add-clique-leaders-option for testing.
gaurav Dec 15, 2025
e68df36
Add a LOGLEVEL to NameRes deployments.
gaurav Dec 18, 2025
d3c8d97
Updated Solr to 9 so we use the latest Solr 9 (currently 9.10).
gaurav Dec 18, 2025
c37c8ae
Updated NameRes Exp to use replace-scaled-clique-identifier-count.
gaurav Dec 18, 2025
fe5034e
Added a Biolink Model Tag.
gaurav Dec 18, 2025
aa74e7b
Reverting to Solr 9.1 to see if that can get it to work.
gaurav Dec 18, 2025
d6d2257
Reverted NCATS Solr to 9.1 as well, just in case.
gaurav Dec 18, 2025
2331d67
Added a Biolink Model tag to NameRes.
gaurav Feb 20, 2026
80725c3
Merge branch 'upgrade-nodenorm-to-v2.4.1' into update-nodenorm-namere…
gaurav Feb 24, 2026
d62e150
Updated renci-exp to fix-ic-for-conflation for testing.
gaurav Feb 26, 2026
0ff967c
Increase NodeNorm info-content DB so accomodate preferred names.
gaurav Jul 15, 2026
9a0d577
Updated NodeNorm Loader to v2.5.0 and 2026jul15.
gaurav Jul 15, 2026
9de1b18
Added Food.txt and updated Gene and Publication max file count.
gaurav Jul 15, 2026
6eb045d
Added a biolink_version to config.json.
gaurav Jul 15, 2026
339405a
Make NodeNorm Redis persistence explicit; disable periodic saves
gaurav Jul 15, 2026
52f0fc1
Incremented Bitnami Redis Helm 23 -> 27
gaurav Jul 15, 2026
c7ee015
Updated Helm charts and dependency.
gaurav Jul 15, 2026
ce36437
Updated Redis version to 8.8.
gaurav Jul 15, 2026
df7c128
Switched nn-loader dev to mode: load.
gaurav Jul 15, 2026
0e74b10
Get the Biolink Version from the root scope.
gaurav Jul 15, 2026
26366fd
Incremented NodeNorm versions to v2.5.1
gaurav Jul 15, 2026
432c5cd
Updated NodeNorm to 2.5.1 and 2026jul15.
gaurav Jul 15, 2026
cb0c8bd
Updated NodeNorm Exp as well.
gaurav Jul 15, 2026
a620ce4
@Claude-generated changes to the Solr load process.
gaurav Jul 16, 2026
3a64779
Reduced memory usage for info-content.
gaurav Jul 22, 2026
1141a21
Updated Babel loader to 2026jul22
gaurav Jul 22, 2026
979fff0
Updated NodeNorm Web to 2026jul22.
gaurav Jul 22, 2026
e4b107c
Relaxed Solr to 9.10
gaurav Jul 23, 2026
42d499a
Updated NameRes to v1.7.0 (Babel 2026jul22).
gaurav Jul 23, 2026
0f03935
Merge branch 'develop' into update-nodenorm-nameres-2025nov4
gaurav Jul 23, 2026
5eea870
Merge branch 'update-nodenorm-nameres-2025nov4' into speed-up-nameres…
gaurav Jul 23, 2026
e856ec5
If this isn't a major change, I don't know what is.
gaurav Jul 23, 2026
10999ed
Reverted to NameRes 1.6.2 for Exp -- so we can test it.
gaurav Jul 23, 2026
94007ce
Fixed Solr tag -- needs to be quoted.
gaurav Jul 23, 2026
57d1ab9
Claude-tweaked memory requirements.
gaurav Jul 23, 2026
bc64f2a
Updated NameRes Exp to v1.7.0.
gaurav Jul 23, 2026
a596b85
Added a stamp to track downloads and changes.
gaurav Jul 23, 2026
3314d72
Added GC logging options for NameRes.
gaurav Jul 24, 2026
d6e3fbb
Claude updates for better performance.
gaurav Jul 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions helm/name-lookup/Chart.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -14,11 +14,11 @@ type: application

# This is the chart version. This version number should be incremented each time you make changes
# to the chart and its templates, including the app version.
version: 0.5.2
version: 0.6.0

# This is the version number of the application being deployed. This version number should be
# incremented each time you make changes to the application.
#
#
# NameRes versions are based on the version of the NameRes Docker image concatenated with the
# Babel release date.
appVersion: 1.5.2_2025sep1
appVersion: 1.7.0_2026jul22
4 changes: 2 additions & 2 deletions helm/name-lookup/ncats-images-meta.yaml
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
nameLookup:
image: ghcr.io/ncatstranslator/nameresolution
version: v1.5.2
version: v1.7.0

solr:
image: solr
version: "9.1"
version: "9.10"

renciPythonImage:
image: ghcr.io/translatorsri/renci-python-image
Expand Down
Binary file modified helm/name-lookup/renci-dev-values-populated.yaml
Binary file not shown.
Binary file modified helm/name-lookup/renci-exp-values-populated.yaml
Binary file not shown.
6 changes: 2 additions & 4 deletions helm/name-lookup/templates/restore-job.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -23,10 +23,8 @@ spec:
requests:
storage: {{ .Values.blocklist.storage }}
{{ end }}
# Because of the way in which the Restore Job works -- by creating fields in
# the Solr database, including copy fields -- it can no longer be re-run on
# failure. Instead, it should fail so we can look at the error output and
# figure out what to do next.
# The restore job now only deletes blocklisted CURIEs (the schema and data
# come baked into the backup), so it is idempotent and safe to re-run.
restartPolicy: Never
containers:
- name: restore-container
Expand Down
203 changes: 42 additions & 161 deletions helm/name-lookup/templates/scripts-config-map.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -9,21 +9,52 @@ data:
#!/bin/bash
set -xa
DATA_DIR="/var/solr/data"
BACKUP_NAME="snapshot.backup"
BACKUP_ZIP="${BACKUP_NAME}.tar.gz"
CORE_DIR="${DATA_DIR}/name_lookup"
BACKUP_ZIP="${DATA_DIR}/snapshot.backup.tar.gz"
BACKUP_URL="{{ .Values.dataUrl }}"
STAMP="${CORE_DIR}/.dataUrl"

rm -rf $DATA_DIR/*
# An empty dataUrl means "leave the volume alone" -- keep whatever core is already
# there (or none). This is the switch for "do not (re)load", same as before.
if [ -z "${BACKUP_URL}" ]; then
echo "dataUrl is empty; leaving any existing core untouched."
exit 0
fi

# The backup is a self-contained Solr core (config + schema + index). Skip the
# download only if the extracted core came from THIS dataUrl: a pod restart with an
# unchanged dataUrl keeps the good core (translator-devops#609), but a new release
# (changed dataUrl) reloads instead of silently serving the old index on a PVC that
# outlived the version bump. A core with no stamp (old logic, or hand-restored)
# reloads once so the served data matches the configured dataUrl.
if [ -f "${CORE_DIR}/core.properties" ] && [ "$(cat "${STAMP}" 2>/dev/null)" = "${BACKUP_URL}" ]; then
echo "Core already present from ${BACKUP_URL}; skipping download."
exit 0
fi

# Otherwise start clean and download.
rm -rf ${DATA_DIR}/*

# Download the file with retries if the download fails or if there is an existing file we can continue.
wget -c --progress=dot:giga --tries=10 --waitretry=10 --timeout=60 --read-timeout=30 -O $DATA_DIR/$BACKUP_ZIP $BACKUP_URL
wget -c --progress=dot:giga --tries=10 --waitretry=10 --timeout=60 --read-timeout=30 -O ${BACKUP_ZIP} ${BACKUP_URL}

cd $DATA_DIR
tar -xf $DATA_DIR/$BACKUP_ZIP -C $DATA_DIR
rm $DATA_DIR/$BACKUP_ZIP
# Extract the core into the Solr home. Solr auto-discovers name_lookup on startup,
# with its schema and config baked in -- nothing else to set up.
tar -xf ${BACKUP_ZIP} -C ${DATA_DIR}
rm ${BACKUP_ZIP}

# Record which backup this core came from, so the next pod start can tell a restart
# (skip) from a version bump (reload). Written after a successful extract so a
# failed download never leaves a stamp that would suppress the retry.
printf '%s' "${BACKUP_URL}" > "${STAMP}"

restore.sh: |-
#!/bin/sh
#
# The backup already contains the loaded index AND the schema/config, so there
# is no collection to create, no replication restore, and no fields to add. The
# only post-restore work is deleting blocklisted CURIEs, if a blocklist is set.
# (This makes the job idempotent and safe to re-run -- see NameResolution#185.)

BLOCKLIST_DIR="/var/blocklist"
BLOCKLIST_CHUNK_SIZE=500
Expand All @@ -46,169 +77,19 @@ data:
COLLECTION_NAME="name_lookup"
SOLR_SERVER=http://{{ include "name-lookup.fullname" . }}-solr-svc:{{ .Values.solr.service.port }}

# liveliness check

HEALTH_ENDPOINT=http://{{ include "name-lookup.fullname" . }}-solr-svc:{{ .Values.solr.service.port }}/solr/admin/cores?action=STATUS
# Wait for Solr and the auto-discovered core to be ready.
HEALTH_ENDPOINT=${SOLR_SERVER}/solr/${COLLECTION_NAME}/admin/ping
response=$(wget --spider --server-response ${HEALTH_ENDPOINT} 2>&1 | grep "HTTP/" | awk '{ print $2 }') >&2
until [ "$response" = "200" ]; do
response=$(wget --spider --server-response ${HEALTH_ENDPOINT} 2>&1 | grep "HTTP/" | awk '{ print $2 }') >&2
echo " -- SOLR is unavailable - sleeping"
echo " -- SOLR is unavailable - sleeping"
sleep 3
done

# solr is ready Now we create collection if it doesn't exist

EXISTS=$(wget -O - ${SOLR_SERVER}/solr/admin/collections?action=LIST | grep name_lookup)

# create collection / shard
if [ -z "$EXISTS" ]
then
wget -O- ${SOLR_SERVER}/solr/admin/collections?action=CREATE'&'name=${COLLECTION_NAME}'&'numShards=1'&'replicationFactor=1
sleep 3
fi

# Setup fields for search
wget --post-data '{"set-user-property": {"update.autoCreateFields": "false"}}' \
--header='Content-Type:application/json' \
-O- ${SOLR_SERVER}/solr/${COLLECTION_NAME}/config
sleep 1

# Restore data
BACKUP_NAME="backup"
CORE_NAME=${COLLECTION_NAME}_shard1_replica_n1
RESTORE_URL=${SOLR_SERVER}/solr/${CORE_NAME}/replication?command=restore'&'location=/var/solr/data/var/solr/data/'&'name=${BACKUP_NAME}
wget -O - $RESTORE_URL
sleep 10
RESTORE_STATUS=$(wget -q -O - ${SOLR_SERVER}/solr/${CORE_NAME}/replication?command=restorestatus 2>&1 | grep "success") >&2
echo "Restore status: " $RESTORE_STATUS
until [ ! -z $RESTORE_STATUS ] ; do
echo "restore not done , probably still loading. Note: if this takes too long please check solr health"
RESTORE_STATUS=$(wget -O - ${SOLR_SERVER}/solr/${CORE_NAME}/replication?command=restorestatus 2>&1 | grep "success") >&2
sleep 10
done
echo "restore done"
wget --post-data '{
"add-field-type" : {
"name": "LowerTextField",
"class": "solr.TextField",
"positionIncrementGap": "100",
"analyzer": {
"tokenizer": {
"class": "solr.StandardTokenizerFactory"
},
"filters": [{
"class": "solr.LowerCaseFilterFactory"
}]
}
}}' \
--header='Content-Type:application/json' \
-O- ${SOLR_SERVER}/solr/${COLLECTION_NAME}/schema
sleep 1
# exactish type taken from https://stackoverflow.com/a/29105025/27310
wget --post-data '{
"add-field-type" : {
"name": "exactish",
"class": "solr.TextField",
"analyzer": {
"tokenizer": {
"class": "solr.KeywordTokenizerFactory"
},
"filters": [{
"class": "solr.LowerCaseFilterFactory"
}]
}
}}' \
--header='Content-Type:application/json' \
-O- ${SOLR_SERVER}/solr/${COLLECTION_NAME}/schema
sleep 1
wget --post-data '{
"add-field": [
{
"name":"names",
"type":"LowerTextField",
"stored": true,
"multiValued": true
},
{
"name":"names_exactish",
"type":"exactish",
"indexed":true,
"stored":true,
"multiValued":true
},
{
"name":"curie",
"type":"string",
"stored":true
},
{
"name": "preferred_name",
"type": "LowerTextField",
"stored": true
},
{
"name": "preferred_name_exactish",
"type": "exactish",
"indexed": true,
"stored": false,
"multiValued": false
},
{
"name": "types",
"type": "string",
"stored": true,
"multiValued": true
},
{
"name": "shortest_name_length",
"type": "pint",
"stored": true
},
{
"name": "curie_suffix",
"type": "plong",
"docValues": true,
"stored": true,
"required": false,
"sortMissingLast": true
},
{
"name": "taxa",
"type": "string",
"stored": true,
"multiValued": true
},
{
"name": "clique_identifier_count",
"type": "pint",
"stored": true
}
]
}' \
--header='Content-Type:application/json' \
-O- ${SOLR_SERVER}/solr/${COLLECTION_NAME}/schema
sleep 1
wget --post-data '{
"add-copy-field" : {
"source": "names",
"dest": "names_exactish"
}}' \
--header='Content-Type:application/json' \
-O- ${SOLR_SERVER}/solr/${COLLECTION_NAME}/schema
wget --post-data '{
"add-copy-field" : {
"source": "preferred_name",
"dest": "preferred_name_exactish"
}}' \
--header='Content-Type:application/json' \
-O- ${SOLR_SERVER}/solr/${COLLECTION_NAME}/schema
sleep 1
echo "Solr core ${COLLECTION_NAME} is up."

# Delete the blocklist terms from the Solr database.
{{ if .Values.blocklist.url }}

BLOCKLIST_DIR=/var/blocklist

# Split the blocklist into files of 500 CURIEs each.
rm -rf ${BLOCKLIST_DIR}/blocklist_*
split -l ${BLOCKLIST_CHUNK_SIZE} ${BLOCKLIST_DIR}/blocklist.txt ${BLOCKLIST_DIR}/blocklist_
Expand Down
9 changes: 9 additions & 0 deletions helm/name-lookup/templates/secrets.yaml
Original file line number Diff line number Diff line change
@@ -1,3 +1,12 @@
{{- /*
Fail fast rather than quietly deploying a blocklist that cannot authenticate: if the
blocklist is switched on (blocklist.url set) there must be a token to go with it. The
token is looked up nil-safely here so that a values file which drops the `secrets` map
entirely still gets this message instead of a raw nil-pointer error.
*/ -}}
{{- if and .Values.blocklist.url (not (.Values.blocklist.secrets).github_personal_access_token) }}
{{- fail "blocklist.url is set but blocklist.secrets.github_personal_access_token is empty. Set the token, or clear blocklist.url to turn the blocklist off. (If the blocklist URL needs no auth, set any non-empty placeholder.)" }}
{{- end }}
{{ if .Values.blocklist.secrets.github_personal_access_token }}
apiVersion: v1
kind: Secret
Expand Down
10 changes: 8 additions & 2 deletions helm/name-lookup/templates/solr-deployment.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -44,9 +44,10 @@ spec:
containers:
- name: {{ .Chart.Name }}
image: "{{ .Values.solr.image.repository }}:{{ .Values.solr.image.tag }}"
# Standalone mode: no -DzkRun/ZooKeeper. The backup is a self-contained
# core that the download init-container extracts into /var/solr/data, so
# Solr auto-discovers name_lookup on startup with its schema baked in.
args:
- '-DzkRun'
- '-DzkClientTimeout=5000'
- '-q'
- '-Dlog4j2.disable.jmx=true'
- '-Dlog4j2.formatMsgNoLookups=true'
Expand All @@ -57,6 +58,11 @@ spec:
value: {{ include "name-lookup.fullname" . }}-solr-svc
- name: GC_TUNE
value: {{ .Values.solr.gc | quote }}
{{- if .Values.solr.gc_log }}
# Overrides Solr's default file-based GC logging (see solr.gc_log in values.yaml).
- name: GC_LOG_OPTS
value: {{ .Values.solr.gc_log | quote }}
{{- end }}
# "-XX:-UseLargePages -XX:+UseG1GC -XX:MaxGCPauseMillis=500 -XX:+UnlockExperimentalVMOptions -XX:G1MaxNewSizePercent=30 -XX:G1NewSizePercent=5 -XX:G1HeapRegionSize=32M -XX:InitiatingHeapOccupancyPercent=70"
# "-XX:-UseLargePages -XX:+UseG1GC -XX:MaxGCPauseMillis=100 -XX:+UnlockExperimentalVMOptions -XX:G1MaxNewSizePercent=20 -XX:G1NewSizePercent=5 -XX:G1HeapRegionSize=32M -XX:InitiatingHeapOccupancyPercent=20"
# "-XX:+UseG1GC -XX:MaxGCPauseMillis=1000 -XX:+UnlockExperimentalVMOptions -XX:G1MaxNewSizePercent=40 -XX:G1NewSizePercent=5 -XX:G1HeapRegionSize=32M -XX:InitiatingHeapOccupancyPercent=90"
Expand Down
6 changes: 6 additions & 0 deletions helm/name-lookup/templates/web-deployment.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -40,11 +40,17 @@ spec:
value: "{{ .Values.app.otel.jaegerHost }}"
- name: "JAEGER_PORT"
value: "{{ .Values.app.otel.jaegerPort }}"
# Logging level.
- name: LOGLEVEL
value: "{{ .Values.app.logLevel }}"
# Babel version information.
- name: BABEL_VERSION
value: "{{ .Values.data.babelVersion }}"
- name: BABEL_VERSION_URL
value: "{{ .Values.data.babelVersionURL }}"
# Biolink Model tag.
- name: BIOLINK_MODEL_TAG
value: "{{ .Values.data.biolinkModelTag }}"
ports:
- name: http
containerPort: {{ .Values.webServer.port }}
Expand Down
Loading