Skip to content

Buffer Base64 Encoding Truncates or Adds Extra Characters #6107

Description

@yunong
  • Version: 0.12-5.x
  • Platform: Mac OS X El Capitan
  • Subsystem: Buffer

Encoding and Decoding a new Buffer object with base64 results in garbled or missing data.

> new Buffer('hello', 'base64').toString('base64');
'hell'
> new Buffer('helloIsItMeYourLookingFor', 'base64').toString('base64');
'helloIsItMeYourLookingFo'
> new Buffer('helloIsItMe', 'base64').toString('base64');
'helloIsItMc='
>

Notice how in the first and second example, the last character is missing, and in the third example, there are extra characters appended to the output.

cc/ @trevnorris

Activity

  1. added
    bufferIssues and PRs related to the buffer subsystem.
    on Apr 7, 2016
  2. kzc commented on Apr 7, 2016

    @kzc

    You're converting ill-formed base64 input to base64. Try this instead:

    Buffer(Buffer('hello', 'utf8').toString('base64'), 'base64').toString();
    
  3. added
    questionIssues asking questions about Node.js.
    on Apr 7, 2016
  4. mscdex commented on Apr 7, 2016

    @mscdex
    Contributor

    @kzc is right, the second argument passed to Buffer() is the encoding for the string passed in as the first argument. 'helloIsItMeYourLookingFor' is not base64-encoded.

  5. added
    invalidIssues and PRs that are invalid.
    and removed
    questionIssues asking questions about Node.js.
    on Apr 7, 2016
  6. davepacheco commented on Apr 8, 2016

    @davepacheco
    Contributor

    Shouldn't invalid input produce an error rather than garbage output?

  7. addaleax commented on Apr 8, 2016

    @addaleax
    Member

    Base64 comes in various forms, the most common using one or two = signs as a form of final padding:

    Buffer.from([0x01, 0x02, 0x03, 0x04]).toString('base64') === 'AQIDBA=='

    This padding makes a buffer encoded with this variant of Base64 always have a length that is a multiple of 4, which is a nice property to have when writing a decoder for it.
    However, it’s not strictly necessary – there is no additional information in those = characters that wouldn’t also be conveyed by the length of the encoded string.

    So, whether you should consider this input invalid is a question you can probably argue about a lot. I’d say there’s no real downside to accepting it, especially since there are some uncommon Base64 encoders which actually leave out the padding; Wikipedia has a quite comprehensive listing of the different variants out there.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bufferIssues and PRs related to the buffer subsystem.invalidIssues and PRs that are invalid.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions