Skip to content

Avoid UTF-8 re-encoding in StringIO character reads - #15956

Merged
josevalim merged 1 commit into
elixir-lang:mainfrom
preciz:optimize-string-io-suffix-scanning
Oct 1, 2026
Merged

josevalim merged 1 commit into
elixir-lang:mainfrom
preciz:optimize-string-io-suffix-scanning

Conversation

@preciz

@preciz preciz commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Assisted-by: Codex CLI:GPT 6

Avoids double work, re-encoding the binary to build the accumulator was not necessary.
x1.2-1.3 speedup, no regression, less memory usage.
Locally I benchmarked with more input types like asian strings, it's all the same outcome.

Bench:

Mix.install([:benchee])

defmodule OldStringIO do
  def read_all(input, count) do
    do_read_all(input, count)
  end

  defp do_read_all("", _count), do: :ok

  defp do_read_all(input, count) do
    case get_chars(input, :unicode, count) do
      {:eof, ""} -> :ok
      {:error, reason} -> raise "Read error: #{inspect(reason)}"
      {_chars, rest} -> do_read_all(rest, count)
    end
  end

  def get_chars("", _encoding, _count) do
    {:eof, ""}
  end

  def get_chars(input, :latin1, count) when byte_size(input) < count do
    {input, ""}
  end

  def get_chars(input, :latin1, count) do
    <<chars::binary-size(^count), rest::binary>> = input
    {chars, rest}
  end

  def get_chars(input, :unicode, count) do
    with {:ok, count} <- split_at(input, count, 0) do
      <<chars::binary-size(^count), rest::binary>> = input
      {chars, rest}
    end
  end

  defp split_at(_, 0, acc),
    do: {:ok, acc}

  defp split_at(<<h::utf8, t::binary>>, count, acc),
    do: split_at(t, count - 1, acc + byte_size(<<h::utf8>>))

  defp split_at(<<_, _::binary>>, _count, _acc),
    do: {:error, :invalid_unicode}

  defp split_at(<<>>, _count, acc),
    do: {:ok, acc}
end

defmodule NewStringIO do
  def read_all(input, count) do
    do_read_all(input, count)
  end

  defp do_read_all("", _count), do: :ok

  defp do_read_all(input, count) do
    case get_chars(input, :unicode, count) do
      {:eof, ""} -> :ok
      {:error, reason} -> raise "Read error: #{inspect(reason)}"
      {_chars, rest} -> do_read_all(rest, count)
    end
  end

  def get_chars("", _encoding, _count) do
    {:eof, ""}
  end

  def get_chars(input, :latin1, count) when byte_size(input) < count do
    {input, ""}
  end

  def get_chars(input, :latin1, count) do
    <<chars::binary-size(^count), rest::binary>> = input
    {chars, rest}
  end

  def get_chars(input, :unicode, count) do
    with {:ok, rest} <- split_at(input, count) do
      size = byte_size(input) - byte_size(rest)
      <<chars::binary-size(^size), _::binary>> = input
      {chars, rest}
    end
  end

  defp split_at(input, 0),
    do: {:ok, input}

  defp split_at(<<_::utf8, rest::binary>>, count),
    do: split_at(rest, count - 1)

  defp split_at(<<_, _::binary>>, _count),
    do: {:error, :invalid_unicode}

  defp split_at(<<>>, _count),
    do: {:ok, <<>>}
end

# Large benchmark inputs (50,000 graphemes each)
ascii_sample = "The quick brown fox jumps over the lazy dog. 1234567890! "

ascii =
  String.duplicate(ascii_sample, ceil(50_000 / String.length(ascii_sample)))
  |> String.slice(0, 50_000)

emoji_sample = "🚀🔥🎉🤖⚡️💡✨🌟🎯🛠️📦🧪"

emoji =
  String.duplicate(emoji_sample, ceil(50_000 / String.length(emoji_sample)))
  |> String.slice(0, 50_000)

# In out, other text types follow the same trend, and 100-character reads
# perform similarly to 1000-character reads. Keep ASCII and emoji representatives.
# Read counts are Unicode code points.
inputs = %{
  "ASCII (1 char)" => {ascii, 1},
  "ASCII (10 chars)" => {ascii, 10},
  "ASCII (1000 chars)" => {ascii, 1000},
  "Emojis (1 char)" => {emoji, 1},
  "Emojis (10 chars)" => {emoji, 10},
  "Emojis (1000 chars)" => {emoji, 1000},
  # Short 10-character reads show a smaller gain for emojis than for ASCII.
  "ASCII short input (10 chars)" => {ascii_sample, 10},
  "Emojis short input (10 chars)" => {emoji_sample, 10}
}

Benchee.run(
  %{
    "main (old)" => fn {input, count} -> OldStringIO.read_all(input, count) end,
    "branch (new)" => fn {input, count} -> NewStringIO.read_all(input, count) end
  },
  inputs: inputs,
  pre_check: :all_same,
  warmup: 1,
  time: 2,
  memory_time: 1
)

Results:

Operating System: Linux
CPU Information: AMD Ryzen 7 8845HS w
Number of Available Cores: 16
Available memory: 54.72 GB
Elixir 1.20.4
Erlang 29.0.5
JIT enabled: true

Benchmark suite executing with the following configuration:
warmup: 1 s
time: 2 s
memory time: 1 s
reduction time: 0 ns
parallel: 1
inputs: ASCII (1 char), ASCII (10 chars), ASCII (1000 chars), ASCII short input (10 chars), Emojis (1 char), Emojis (10 chars), Emojis (1000 chars), Emojis short input (10 chars)
Estimated total run time: 1 min 4 s
Excluding outliers: false

##### With input ASCII (1 char) #####
Name                   ips        average  deviation         median         99th %
branch (new)        633.09        1.58 ms     ±6.04%        1.55 ms        2.08 ms
main (old)          514.57        1.94 ms     ±5.83%        1.92 ms        2.57 ms

Comparison: 
branch (new)        633.09
main (old)          514.57 - 1.23x slower +0.36 ms

Memory usage statistics:

Name            Memory usage
branch (new)         9.16 MB
main (old)          12.21 MB - 1.33x memory usage +3.05 MB

**All measurements for memory usage were the same**

##### With input ASCII (10 chars) #####
Name                   ips        average  deviation         median         99th %
branch (new)        1.72 K      582.97 μs     ±5.73%      578.14 μs      698.50 μs
main (old)          1.17 K      857.58 μs     ±6.87%      842.86 μs     1171.09 μs

Comparison: 
branch (new)        1.72 K
main (old)          1.17 K - 1.47x slower +274.61 μs

Memory usage statistics:

Name            Memory usage
branch (new)         4.39 MB
main (old)           5.72 MB - 1.30x memory usage +1.34 MB

**All measurements for memory usage were the same**

##### With input ASCII (1000 chars) #####
Name                   ips        average  deviation         median         99th %
branch (new)        2.20 K      454.15 μs     ±6.79%      445.32 μs      573.41 μs
main (old)          1.37 K      727.47 μs     ±4.55%      723.15 μs      869.56 μs

Comparison: 
branch (new)        2.20 K
main (old)          1.37 K - 1.60x slower +273.32 μs

Memory usage statistics:

Name            Memory usage
branch (new)         3.82 MB
main (old)           4.97 MB - 1.30x memory usage +1.15 MB

**All measurements for memory usage were the same**

##### With input ASCII short input (10 chars) #####
Name                   ips        average  deviation         median         99th %
branch (new)        1.28 M        0.78 μs  ±1045.11%        0.70 μs        1.30 μs
main (old)          0.90 M        1.11 μs   ±762.58%        1.02 μs        1.86 μs

Comparison: 
branch (new)        1.28 M
main (old)          0.90 M - 1.42x slower +0.33 μs

Memory usage statistics:

Name            Memory usage
branch (new)         5.58 KB
main (old)           7.18 KB - 1.29x memory usage +1.60 KB

**All measurements for memory usage were the same**

##### With input Emojis (1 char) #####
Name                   ips        average  deviation         median         99th %
branch (new)        474.51        2.11 ms     ±6.30%        2.08 ms        2.68 ms
main (old)          397.00        2.52 ms     ±5.11%        2.46 ms        3.21 ms

Comparison: 
branch (new)        474.51
main (old)          397.00 - 1.20x slower +0.41 ms

Memory usage statistics:

Name            Memory usage
branch (new)        10.68 MB
main (old)          14.24 MB - 1.33x memory usage +3.56 MB

**All measurements for memory usage were the same**

##### With input Emojis (10 chars) #####
Name                   ips        average  deviation         median         99th %
branch (new)        1.21 K        0.83 ms     ±8.82%        0.82 ms        1.16 ms
main (old)          0.83 K        1.20 ms     ±5.90%        1.18 ms        1.52 ms

Comparison: 
branch (new)        1.21 K
main (old)          0.83 K - 1.45x slower +0.37 ms

Memory usage statistics:

Name            Memory usage
branch (new)         5.25 MB
main (old)           6.81 MB - 1.30x memory usage +1.56 MB

**All measurements for memory usage were the same**

##### With input Emojis (1000 chars) #####
Name                   ips        average  deviation         median         99th %
branch (new)        1.45 K        0.69 ms     ±6.87%        0.68 ms        0.83 ms
main (old)          0.97 K        1.03 ms     ±7.96%        1.02 ms        1.21 ms

Comparison: 
branch (new)        1.45 K
main (old)          0.97 K - 1.49x slower +0.34 ms

Memory usage statistics:

Name            Memory usage
branch (new)         4.46 MB
main (old)           5.80 MB - 1.30x memory usage +1.34 MB

**All measurements for memory usage were the same**

##### With input Emojis short input (10 chars) #####
Name                   ips        average  deviation         median         99th %
branch (new)        3.18 M      314.71 ns  ±2241.62%         271 ns         441 ns
main (old)          2.67 M      373.90 ns  ±1465.20%         341 ns         611 ns

Comparison: 
branch (new)        3.18 M
main (old)          2.67 M - 1.19x slower +59.19 ns

Memory usage statistics:

Name            Memory usage
branch (new)         1.41 KB
main (old)           1.80 KB - 1.28x memory usage +0.40 KB

**All measurements for memory usage were the same**

@josevalim
josevalim merged commit 41da560 into elixir-lang:main Oct 1, 2026
15 checks passed
@josevalim

Copy link
Copy Markdown
Member

💚 💙 💜 💛 ❤️

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants