Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion .jules/bolt.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,10 @@
## 2024-05-18 - Language Detector Optimization
**Learning:** `LanguageDetector.detect` uses `Regex("[^\\p{L}]+")` inside `detectByHeuristics`, creating a new Regex instance on every call. This method is called via `ProofreadHelper` which is likely triggered on user typing or interaction. Pre-compiling the Regex as a top-level constant avoids the overhead of regex compilation for every detection call.
**Action:** Always check for repeated Regex instantiations in frequently called string processing or detection methods and extract them to top-level or companion object properties.

## 2024-05-18 - StringUtils Whitespace Split Optimization
**Learning:** `StringUtils.splitOnWhitespace` creates a `Regex("\\s+")` inline every time it is called. Since this is an extension function used across settings and text processing, this causes unnecessary allocations.
**Action:** Abstract frequent inline regex usages into `private val` properties at the object or file level to avoid repeated regex compilation.

## 2024-05-18 - Pre-compile Regex in hot dictionary loops
**Learning:** Instantiating `Regex` objects inside dictionary processing loops (like adding words in `UserBinaryDictionary` or `AppsBinaryDictionary`) causes unnecessary allocations and compilation overhead on every item iteration.
**Action:** Always pre-compile `Regex` objects as `private val` properties in the class or companion object to eliminate per-iteration compilation and allocation pressure.
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,7 @@ class AppsBinaryDictionary private constructor(
private const val TAG = "AppsBinaryDictionary"
private const val NAME = "apps"

private val SPACE_REGEX = Regex("\\s+")
private const val FREQUENCY_FOR_APPS = 100
private const val FREQUENCY_FOR_APPS_BIGRAM = 200

Expand Down Expand Up @@ -75,7 +76,7 @@ class AppsBinaryDictionary private constructor(
var ngramContext = NgramContext.getEmptyPrevWordsContext(
BinaryDictionary.MAX_PREV_WORD_COUNT_FOR_N_GRAM
)
for (word in appLabel.split(Regex("\\s+"))) {
for (word in appLabel.split(SPACE_REGEX)) {
if (word.isEmpty()) continue
if (DEBUG_DUMP) {
Log.d(TAG, "addName word = $word")
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -79,6 +79,7 @@ class UserBinaryDictionary protected constructor(
Words.FREQUENCY
)

private val SPACE_REGEX = Regex("\\s+")
private const val NAME = "userunigram"

fun getDictionary(
Expand Down Expand Up @@ -195,7 +196,7 @@ class UserBinaryDictionary protected constructor(
)
}
// ponytail: split phrase into unigrams and n-grams for next-word prediction
val parts = word.split(Regex("\\s+"))
val parts = word.split(SPACE_REGEX)
if (parts.size > 1) {
for (part in parts) {
if (part.length <= MAX_WORD_LENGTH && part.isNotEmpty() && part != word) {
Expand Down
Loading