Byte & Character Counter
Words, Lines, UTF-8
Count UTF-8 bytes, characters, words, lines, and sentences instantly. Free online byte counter and string size tool.
This conversion runs only in your browser — nothing is uploaded.
Counts run locally in your browser — nothing is uploaded.
Characters
0
Words
0
Lines
0
Bytes (UTF-8)
0
- Characters (no spaces)0
- Sentences0
Input
What is this tool?
This string counter measures text length from several practical angles: characters, characters without whitespace, words, lines, sentences, and UTF-8 bytes. Counts update live as you type; nothing is uploaded.
Different limits count different things. A UI field may cap characters, a database column may cap bytes, an editorial brief may cap words, and an API may care about encoded payload size. Non-ASCII characters can use multiple UTF-8 bytes even when they look like one glyph.
Common use cases
- Check meta descriptions, social posts, or form fields against character vs byte limits
- Compare translation length before shipping UI copy
- Measure API or log payload size in UTF-8 bytes before sending
How to use
- Paste or type text into the input. KPIs (characters, words, lines, UTF-8 bytes) update live at the top.
- Use characters without spaces and sentences in the chip row for finer checks.
- Load Sample to see ASCII vs Hangul/emoji byte differences.
- Copy the stats summary or Clear when done.
Examples
| Input / setting | Output | Notes |
|---|---|---|
Hello world. | chars 12 · no-spaces 11 · words 2 · lines 1 · sentences 1 · UTF-8 12 | Simple ASCII. |
안녕하세요 | chars 5 · words 1 · UTF-8 15 | Hangul: 3 bytes per syllable in UTF-8. |
Hi 👋 | chars 5 · words 2 · UTF-8 7 | Space + waving hand: JS length 5 (👋 = 2 UTF-16 units), UTF-8 7. |
One.\nTwo! | lines 2 · sentences 2 · words 2 | Newline starts a new line; ! ends a sentence. |
Practical pitfalls
- UTF-8 bytes ≠ characters: Hangul, CJK, and emoji inflate byte size; do not use
string.lengthalone for storage quotas. - Words are whitespace-based: CJK or glued tokens without spaces usually count as one word—not a linguistic tokenizer.
- Character count is JavaScript UTF-16 length: some emoji are two “characters” even though they render as one glyph.
- Sentence detection is punctuation-heuristic;
U.S.A.-style abbreviations can over-count.
References
이 도구는 무엇인가요?
문자·공백 제외 문자·단어·줄·문장·UTF-8 바이트를 한곳에서 셉니다. 입력하면 KPI가 바로 갱신되며, 데이터는 업로드되지 않습니다.
UI는 글자 수, DB는 바이트, 원고는 단어, API는 인코딩 크기를 제한하는 경우가 많습니다. 한글·이모지는 보이는 글자보다 UTF-8 바이트가 클 수 있습니다.
자주 쓰는 경우
- 메타 설명·SNS·입력란의 글자/바이트 한도 확인
- UI 카피 번역 길이 비교
- API·로그 페이로드의 UTF-8 바이트 크기 점검
사용 방법
- 텍스트를 붙여넣습니다. 상단 KPI(문자·단어·줄·UTF-8 바이트)가 즉시 갱신됩니다.
- 칩 행에서 공백 제외 문자·문장 수를 확인합니다.
- Sample로 ASCII와 한글·이모지 바이트 차이를 봅니다.
- 통계 요약을 복사하거나 Clear로 비웁니다.
예시
| 입력 / 설정 | 출력 | 참고 |
|---|---|---|
Hello world. | chars 12 · no-spaces 11 · words 2 · lines 1 · sentences 1 · UTF-8 12 | 단순 ASCII. |
안녕하세요 | chars 5 · words 1 · UTF-8 15 | 한글 음절당 UTF-8 3바이트. |
Hi 👋 | chars 5 · words 2 · UTF-8 7 | 👋는 UTF-16 2유닛·UTF-8 4바이트(전체 7). |
One.\nTwo! | lines 2 · sentences 2 · words 2 | 줄바꿈·문장 부호 반영. |
실무에서 흔한 실수
- UTF-8 바이트 ≠ 글자 수: 한글·CJK·이모지는 저장 한도에
length만 쓰면 안 됩니다. - 단어는 공백 기준: 띄어쓰기 없는 한글 문장은 단어 1개로 잡힐 수 있습니다(형태소 분석 아님).
- 글자 수는 JS UTF-16 길이입니다. 일부 이모지는 화면상 1글자여도 2로 셉니다.
- 문장은
.!?휴리스틱이라 약어에서 과다 집계될 수 있습니다.
참고자료
このツールについて
文字数・空白除く文字数・単語・行・文・UTF-8 バイトをまとめて数えます。入力すると KPI が即更新され、アップロードはありません。
UI は文字、DB はバイト、原稿は単語、API はエンコードサイズを制限することがあります。日本語・絵文字は見た目より UTF-8 バイトが大きくなりがちです。
よくある用途
- メタ説明・SNS・入力欄の文字/バイト上限チェック
- UI コピー翻訳の長さ比較
- API・ログペイロードの UTF-8 バイト確認
使い方
- テキストを貼り付けます。上部 KPI(文字・単語・行・UTF-8 バイト)がすぐ更新されます。
- チップ行で空白除く文字数・文数を確認します。
- Sample で ASCII と日本語・絵文字のバイト差を見ます。
- 統計サマリーをコピーするか Clear します。
例
| 入力 / 設定 | 出力 | メモ |
|---|---|---|
Hello world. | chars 12 · no-spaces 11 · words 2 · lines 1 · sentences 1 · UTF-8 12 | 単純 ASCII。 |
こんにちは | chars 5 · words 1 · UTF-8 15 | ひらがなは UTF-8 で多バイト。 |
Hi 👋 | chars 5 · words 2 · UTF-8 7 | 👋 は UTF-16 で 2、UTF-8 で 4(合計 7)。 |
One.\nTwo! | lines 2 · sentences 2 · words 2 | 改行と文末記号を反映。 |
実務での落とし穴
- UTF-8 バイト ≠ 文字数: 日本語・CJK・絵文字では保存上限に
lengthだけでは足りません。 - 単語は空白区切り: 空白のない日本語は 単語 1 になりやすい(形態素解析ではありません)。
- 文字数は JS の UTF-16 長です。一部の絵文字は見た目 1 でも 2 と数えます。
- 文は
.!?ヒューリスティックで、略語で過大になることがあります。
参考資料
¿Qué es esta herramienta?
Mide caracteres, caracteres sin espacios, palabras, líneas, oraciones y bytes UTF-8. Los KPI se actualizan al escribir; no se sube nada.
Distintos límites cuentan cosas distintas: UI por caracteres, BD por bytes, editorial por palabras, API por tamaño codificado. El texto no ASCII suele ocupar varios bytes UTF-8 por glifo.
Casos de uso comunes
- Comprobar meta, redes o campos frente a límites de caracteres vs bytes
- Comparar longitud de traducciones de UI
- Medir el tamaño UTF-8 de payloads de API o logs
Cómo usar
- Pega o escribe texto. Los KPI (caracteres, palabras, líneas, bytes UTF-8) se actualizan arriba.
- Revisa caracteres sin espacios y oraciones en la fila de chips.
- Usa Sample para ver ASCII frente a Hangul/emoji.
- Copia el resumen o Clear al terminar.
Ejemplos
| Entrada / configuración | Salida | Notas |
|---|---|---|
Hello world. | chars 12 · no-spaces 11 · words 2 · lines 1 · sentences 1 · UTF-8 12 | ASCII simple. |
안녕하세요 | chars 5 · words 1 · UTF-8 15 | Hangul: ~3 bytes UTF-8 por sílaba. |
Hi 👋 | chars 5 · words 2 · UTF-8 7 | 👋 = 2 unidades UTF-16, 4 bytes UTF-8 (total 7). |
One.\nTwo! | lines 2 · sentences 2 · words 2 | Nueva línea y puntuación. |
Errores prácticos
- Bytes UTF-8 ≠ caracteres: Hangul, CJK y emoji inflan el tamaño; no uses solo
lengthpara cuotas de almacenamiento. - Palabras por espacios: texto CJK sin espacios suele contar como una palabra—no es un tokenizador lingüístico.
- El conteo de caracteres es longitud UTF-16 de JS: algunos emoji son dos “caracteres”.
- Las oraciones son heurísticas con
.!?; abreviaturas pueden sobrercontar.
Referencias
Qu'est-ce que cet outil ?
Mesure caractères, caractères sans espaces, mots, lignes, phrases et octets UTF-8. Les KPI se mettent à jour en direct ; rien n’est téléversé.
Les limites ne comptent pas la même chose : UI en caractères, base en octets, brief en mots, API en taille encodée. Le non-ASCII prend souvent plusieurs octets UTF-8 par glyphe.
Cas d'usage courants
- Vérifier méta, posts ou champs face aux plafonds caractères vs octets
- Comparer la longueur des traductions d’UI
- Mesurer la taille UTF-8 des payloads API ou logs
Comment utiliser
- Collez ou saisissez du texte. Les KPI (caractères, mots, lignes, octets UTF-8) se mettent à jour en haut.
- Consultez caractères sans espaces et phrases dans la rangée de chips.
- Utilisez Sample pour comparer ASCII et Hangul/emoji.
- Copiez le résumé ou Clear.
Exemples
| Entrée / réglage | Sortie | Notes |
|---|---|---|
Hello world. | chars 12 · no-spaces 11 · words 2 · lines 1 · sentences 1 · UTF-8 12 | ASCII simple. |
안녕하세요 | chars 5 · words 1 · UTF-8 15 | Hangul : ~3 octets UTF-8 par syllabe. |
Hi 👋 | chars 5 · words 2 · UTF-8 7 | 👋 = 2 unités UTF-16, 4 octets UTF-8 (total 7). |
One.\nTwo! | lines 2 · sentences 2 · words 2 | Nouvelle ligne et ponctuation. |
Pièges courants
- Octets UTF-8 ≠ caractères : Hangul, CJK et emoji gonflent la taille ; ne vous fiez pas au seul
lengthpour les quotas. - Mots découpés sur espaces : un texte CJK sans espaces compte souvent comme un mot—pas de tokeniseur linguistique.
- Le compte de caractères est la longueur UTF-16 JS : certains emoji font deux « caractères ».
- Les phrases sont heuristiques (
.!?) ; les abréviations peuvent surcompter.
Références
这个工具是什么?
统计字符、去空白字符、词、行、句子和 UTF-8 字节。输入时顶部 KPI 即时更新,不会上传数据。
不同限制统计不同维度:UI 常限字符、数据库常限字节、文案常限词数、API 关心编码后体积。非 ASCII 往往一个字形对应多个 UTF-8 字节。
常见用途
- 检查 meta、社交文案或表单的字符/字节上限
- 对比 UI 文案翻译长度
- 测量 API 或日志 payload 的 UTF-8 字节大小
如何使用
- 粘贴或输入文本。顶部 KPI(字符、词、行、UTF-8 字节)即时更新。
- 在芯片行查看去空格字符数与句子数。
- 用 Sample 对比 ASCII 与韩文/emoji 的字节差异。
- 复制统计摘要或 Clear 清空。
示例
| 输入 / 设置 | 输出 | 说明 |
|---|---|---|
Hello world. | chars 12 · no-spaces 11 · words 2 · lines 1 · sentences 1 · UTF-8 12 | 简单 ASCII。 |
안녕하세요 | chars 5 · words 1 · UTF-8 15 | 韩文音节多为 3 字节 UTF-8。 |
Hi 👋 | chars 5 · words 2 · UTF-8 7 | 👋 为 2 个 UTF-16 单元、4 个 UTF-8 字节(合计 7)。 |
One.\nTwo! | lines 2 · sentences 2 · words 2 | 换行与句末标点。 |
实践中的常见误区
- UTF-8 字节 ≠ 字符数:韩文、CJK、emoji 会抬高字节;不要只用
length估存储配额。 - 词数按空白分割:无空格的中文通常计为 1 个词——不是分词器。
- 字符数是 JS 的 UTF-16 长度:部分 emoji 显示为 1 个字形却计为 2。
- 句子以
.!?启发式检测,缩写可能多计。
参考资料
這個工具是什麼?
統計字元、去空白字元、詞、行、句子與 UTF-8 位元組。輸入時頂部 KPI 即時更新,不會上傳資料。
不同限制統計不同維度:UI 常限字元、資料庫常限位元組、文案常限詞數、API 關心編碼後體積。非 ASCII 往往一個字形對應多個 UTF-8 位元組。
常見用途
- 檢查 meta、社群文案或表單的字元/位元組上限
- 對比 UI 文案翻譯長度
- 測量 API 或日誌 payload 的 UTF-8 位元組大小
如何使用
- 貼上或輸入文字。頂部 KPI(字元、詞、行、UTF-8 位元組)即時更新。
- 在晶片列查看去空白字元數與句子數。
- 用 Sample 對比 ASCII 與韓文/emoji 的位元組差異。
- 複製統計摘要或 Clear 清空。
範例
| 輸入 / 設定 | 輸出 | 說明 |
|---|---|---|
Hello world. | chars 12 · no-spaces 11 · words 2 · lines 1 · sentences 1 · UTF-8 12 | 簡單 ASCII。 |
안녕하세요 | chars 5 · words 1 · UTF-8 15 | 韓文音節多為 3 位元組 UTF-8。 |
Hi 👋 | chars 5 · words 2 · UTF-8 7 | 👋 為 2 個 UTF-16 單元、4 個 UTF-8 位元組(合計 7)。 |
One.\nTwo! | lines 2 · sentences 2 · words 2 | 換行與句末標點。 |
實務上常見陷阱
- UTF-8 位元組 ≠ 字元數:韓文、CJK、emoji 會抬高位元組;不要只用
length估儲存配額。 - 詞數依空白分割:無空白的中文通常計為 1 個詞——不是斷詞器。
- 字元數是 JS 的 UTF-16 長度:部分 emoji 顯示為 1 個字形卻計為 2。
- 句子以
.!?啟發式偵測,縮寫可能多計。
參考資料
Last reviewed: 2026-07-27
Frequently asked questions
- Why do UTF-8 bytes differ from character count?
- Character count uses JavaScript string length (UTF-16 code units). UTF-8 byte size is what TextEncoder produces. Non-ASCII text (Hangul, emoji, accents) often needs more than one byte per visible character—and some emoji use two UTF-16 units but four UTF-8 bytes.
- How are words counted?
- After trimming, the text is split on whitespace (spaces, tabs, newlines). Consecutive spaces still count as one break. Text without spaces—such as many CJK phrases—counts as one word.
- Are sentences counted accurately for every language?
- Sentences are heuristic: runs ending with `.`, `!`, or `?`. Abbreviations (e.g. `Dr.`) and scripts that do not use those marks can skew the count.
- Is text uploaded?
- No. All counts run locally in your browser.