fix(bedrock-kb-retrieval): preserve non-ASCII characters in tool output - #4372
Open
Baddala-Govardhan wants to merge 1 commit into
Open
fix(bedrock-kb-retrieval): preserve non-ASCII characters in tool output#4372Baddala-Govardhan wants to merge 1 commit into
Baddala-Govardhan wants to merge 1 commit into
Conversation
QueryKnowledgeBases and ListKnowledgeBases serialized results with json.dumps() using its default ensure_ascii=True, turning CJK text, accented characters, and emoji into \uXXXX escape sequences instead of readable text. This inflated token usage and made raw tool output unreadable. Pass ensure_ascii=False at both serialization points since the transport is already UTF-8. Fixes awslabs#3820
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #3820
json.dumps()defaults toensure_ascii=True, so any non-ASCII text in aKnowledge Base (Japanese/Korean/Chinese, emoji, accented characters) was
coming back as
\uXXXXescapes instead of the actual text. Besides beingunreadable, it also bloats token usage since each CJK char turns into a
6-char escape sequence.
Added
ensure_ascii=Falsein the two places that serialize results(
server.pyandknowledgebases/retrieval.py) since the transport isUTF-8 anyway.
Added tests for both spots to make sure non-ASCII text round-trips without
getting escaped. Ran the full suite + pre-commit, all green.
By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of the project license.