{"thread":{"id":"45770","subject":"Git archive doesn't fully support zip64","startedAt":"2017-04-21T21:08:37Z","lastAt":"2017-05-01T08:30:33Z","messageCount":44,"participants":["Keith Goldfarb","Peter Krefting","Johannes Sixt","René Scharfe","Junio C Hamano","Torsten Bögershausen"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"317520","messageId":"3C736801-6BB8-41CC-88FF-C42FC853A736@blackthorn-media.com","threadId":"45770","inReplyTo":null,"subject":"Git archive doesn't fully support zip64","fromName":"Keith Goldfarb","fromEmail":"keith@blackthorn-media.com","sentAt":"2017-04-21T21:08:28Z","receivedAt":"2017-04-21T21:08:37Z","isPatch":false,"sender":{"key":"keith@blackthorn-media.com","avatar":null},"body":"Dear git,\n\ngit archive, when writing a zip file, has a silent 4GB file size limit (on the inputs as well as the output), as it doesn’t fully support zip64.\n\nAlthough a zip archive written by git which is larger than 4GB can be often still be unzipped, it won’t be fully useable and some tools (e.g. zipdetails) can’t read it at all.\n\nI suggest either git should be changed to fully support zip64 or it should give a proper error when either an input or the output is too large.\n\nThanks,\n\nK.\n\n\n"},{"id":"317548","messageId":"alpine.DEB.2.11.1704222019420.12779@perkele.intern.softwolves.pp.se","threadId":"45770","inReplyTo":"3C736801-6BB8-41CC-88FF-C42FC853A736@blackthorn-media.com","subject":"[PATCH] archive-zip: Add zip64 headers when file size is too large for 32 bits","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2017-04-22T19:22:37Z","receivedAt":"2017-04-22T19:32:21Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"If the size of the files in the archive cannot be expressed in 32 bits, or\nthe offset in the zip file itself, add zip64 local headers with the actual\nsize. If we do find such entries, we also set a flag to force the creation\nof a zip64 end of central directory record.\n\nSigned-off-by: Peter Krefting <peter@softwolves.pp.se>\n---\n  archive-zip.c | 50 +++++++++++++++++++++++++++++++++++++++++---------\n  1 file changed, 41 insertions(+), 9 deletions(-)\n\n> git archive, when writing a zip file, has a silent 4GB file size \n> limit (on the inputs as well as the output), as it doesn’t fully \n> support zip64.\n\nYeah, it seems that the zip64 support that was added was to support \nmore than 65535 files, but it did not add support for 64-bit file \nsizes or ZIP archives.\n\nTry the below patch, it seems to work for me with a repository with \ntwo files of 4 Gbyte plus a few bytes. Haven't tested the case where \nthe archive itself is larger than 4 Gbyte, but that ought to work, too.\n\ndiff --git a/archive-zip.c b/archive-zip.c\nindex b429a8d..c76a9b4 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -10,6 +10,7 @@\n\n  static int zip_date;\n  static int zip_time;\n+static int zip_zip64;\n\n  static unsigned char *zip_dir;\n  static unsigned int zip_dir_size;\n@@ -88,6 +89,16 @@ struct zip_extra_mtime {\n  \tunsigned char _end[1];\n  };\n\n+struct zip_extra_zip64 {\n+\tunsigned char magic[2];\n+\tunsigned char extra_size[2];\n+\tunsigned char size[8];\n+\tunsigned char compressed_size[8];\n+\tunsigned char offset[8];\n+\tunsigned char disk[4];\n+\tunsigned char _end[1];\n+};\n+\n  struct zip64_dir_trailer {\n  \tunsigned char magic[4];\n  \tunsigned char record_size[8];\n@@ -122,6 +133,9 @@ struct zip64_dir_trailer_locator {\n  #define ZIP_EXTRA_MTIME_SIZE\toffsetof(struct zip_extra_mtime, _end)\n  #define ZIP_EXTRA_MTIME_PAYLOAD_SIZE \\\n  \t(ZIP_EXTRA_MTIME_SIZE - offsetof(struct zip_extra_mtime, flags))\n+#define ZIP_EXTRA_ZIP64_SIZE\toffsetof(struct zip_extra_zip64, _end)\n+#define ZIP_EXTRA_ZIP64_PAYLOAD_SIZE \\\n+\t(ZIP_EXTRA_ZIP64_SIZE - offsetof(struct zip_extra_zip64, size))\n  #define ZIP64_DIR_TRAILER_SIZE\toffsetof(struct zip64_dir_trailer, _end)\n  #define ZIP64_DIR_TRAILER_RECORD_SIZE \\\n  \t(ZIP64_DIR_TRAILER_SIZE - \\\n@@ -219,19 +233,25 @@ static void set_zip_dir_data_desc(struct zip_dir_header *header,\n  \t\t\t\t  unsigned long compressed_size,\n  \t\t\t\t  unsigned long crc)\n  {\n+\tint clamped = 0;\n  \tcopy_le32(header->crc32, crc);\n-\tcopy_le32(header->compressed_size, compressed_size);\n-\tcopy_le32(header->size, size);\n+\tcopy_le32(header->compressed_size, clamp_max(compressed_size, 0xFFFFFFFFU, &clamped));\n+\tcopy_le32(header->size, clamp_max(size, 0xFFFFFFFFU, &clamped));\n+\tif (clamped)\n+\t\tzip_zip64 = 1;\n  }\n\n  static void set_zip_header_data_desc(struct zip_local_header *header,\n  \t\t\t\t     unsigned long size,\n  \t\t\t\t     unsigned long compressed_size,\n-\t\t\t\t     unsigned long crc)\n+\t\t\t\t     unsigned long crc,\n+\t\t\t\t     int *clamped)\n  {\n  \tcopy_le32(header->crc32, crc);\n-\tcopy_le32(header->compressed_size, compressed_size);\n-\tcopy_le32(header->size, size);\n+\tcopy_le32(header->compressed_size, clamp_max(compressed_size, 0xFFFFFFFFU, clamped));\n+\tcopy_le32(header->size, clamp_max(size, 0xFFFFFFFFU, clamped));\n+\tif (clamped)\n+\t\tzip_zip64 = 1;\n  }\n\n  static int has_only_ascii(const char *s)\n@@ -279,6 +299,7 @@ static int write_zip_entry(struct archiver_args *args,\n  \tint is_binary = -1;\n  \tconst char *path_without_prefix = path + args->baselen;\n  \tunsigned int creator_version = 0;\n+\tint clamped = 0;\n\n  \tcrc = crc32(0, NULL, 0);\n\n@@ -376,7 +397,7 @@ static int write_zip_entry(struct archiver_args *args,\n  \tcopy_le16(dirent.comment_length, 0);\n  \tcopy_le16(dirent.disk, 0);\n  \tcopy_le32(dirent.attr2, attr2);\n-\tcopy_le32(dirent.offset, zip_offset);\n+\tcopy_le32(dirent.offset, clamp_max(zip_offset, 0xFFFFFFFFU, &clamped));\n\n  \tcopy_le32(header.magic, 0x04034b50);\n  \tcopy_le16(header.version, 10);\n@@ -384,15 +405,26 @@ static int write_zip_entry(struct archiver_args *args,\n  \tcopy_le16(header.compression_method, method);\n  \tcopy_le16(header.mtime, zip_time);\n  \tcopy_le16(header.mdate, zip_date);\n-\tset_zip_header_data_desc(&header, size, compressed_size, crc);\n+\tset_zip_header_data_desc(&header, size, compressed_size, crc, &clamped);\n  \tcopy_le16(header.filename_length, pathlen);\n-\tcopy_le16(header.extra_length, ZIP_EXTRA_MTIME_SIZE);\n+\tcopy_le16(header.extra_length, ZIP_EXTRA_MTIME_SIZE + (clamped ? ZIP_EXTRA_ZIP64_SIZE : 0));\n  \twrite_or_die(1, &header, ZIP_LOCAL_HEADER_SIZE);\n  \tzip_offset += ZIP_LOCAL_HEADER_SIZE;\n  \twrite_or_die(1, path, pathlen);\n  \tzip_offset += pathlen;\n  \twrite_or_die(1, &extra, ZIP_EXTRA_MTIME_SIZE);\n  \tzip_offset += ZIP_EXTRA_MTIME_SIZE;\n+\tif (clamped) {\n+\t\tstruct zip_extra_zip64 extra_zip64;\n+\t\tcopy_le16(extra_zip64.magic, 0x0001);\n+\t\tcopy_le16(extra_zip64.extra_size, ZIP_EXTRA_ZIP64_PAYLOAD_SIZE);\n+\t\tcopy_le64(extra_zip64.size, size);\n+\t\tcopy_le64(extra_zip64.compressed_size, compressed_size);\n+\t\tcopy_le64(extra_zip64.offset, zip_offset);\n+\t\tcopy_le32(extra_zip64.disk, 0);\n+\t\twrite_or_die(1, &extra_zip64, ZIP_EXTRA_ZIP64_SIZE);\n+\t\tzip_offset += ZIP_EXTRA_ZIP64_SIZE;\n+\t}\n  \tif (stream && method == 0) {\n  \t\tunsigned char buf[STREAM_BUFFER_SIZE];\n  \t\tssize_t readlen;\n@@ -538,7 +570,7 @@ static void write_zip_trailer(const unsigned char *sha1)\n  \tcopy_le16(trailer.comment_length, sha1 ? GIT_SHA1_HEXSZ : 0);\n\n  \twrite_or_die(1, zip_dir, zip_dir_offset);\n-\tif (clamped)\n+\tif (clamped || zip_zip64)\n  \t\twrite_zip64_trailer();\n  \twrite_or_die(1, &trailer, ZIP_DIR_TRAILER_SIZE);\n  \tif (sha1)\n-- \n2.1.4\n"},{"id":"317550","messageId":"37eb7c14-eb61-7a63-bdf0-ee1ccf40723f@kdbg.org","threadId":"45770","inReplyTo":"alpine.DEB.2.11.1704222019420.12779@perkele.intern.softwolves.pp.se","subject":"Re: [PATCH] archive-zip: Add zip64 headers when file size is too large for 32 bits","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2017-04-22T21:52:48Z","receivedAt":"2017-04-22T21:52:58Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 22.04.2017 um 21:22 schrieb Peter Krefting:\n> @@ -279,6 +299,7 @@ static int write_zip_entry(struct archiver_args *args,\n>      int is_binary = -1;\n>      const char *path_without_prefix = path + args->baselen;\n>      unsigned int creator_version = 0;\n> +    int clamped = 0;\n>\n>      crc = crc32(0, NULL, 0);\n>\n> @@ -376,7 +397,7 @@ static int write_zip_entry(struct archiver_args *args,\n>      copy_le16(dirent.comment_length, 0);\n>      copy_le16(dirent.disk, 0);\n>      copy_le32(dirent.attr2, attr2);\n> -    copy_le32(dirent.offset, zip_offset);\n> +    copy_le32(dirent.offset, clamp_max(zip_offset, 0xFFFFFFFFU,\n> &clamped));\n>\n>      copy_le32(header.magic, 0x04034b50);\n>      copy_le16(header.version, 10);\n> @@ -384,15 +405,26 @@ static int write_zip_entry(struct archiver_args\n> *args,\n>      copy_le16(header.compression_method, method);\n>      copy_le16(header.mtime, zip_time);\n>      copy_le16(header.mdate, zip_date);\n> -    set_zip_header_data_desc(&header, size, compressed_size, crc);\n> +    set_zip_header_data_desc(&header, size, compressed_size, crc,\n> &clamped);\n>      copy_le16(header.filename_length, pathlen);\n> -    copy_le16(header.extra_length, ZIP_EXTRA_MTIME_SIZE);\n> +    copy_le16(header.extra_length, ZIP_EXTRA_MTIME_SIZE + (clamped ?\n> ZIP_EXTRA_ZIP64_SIZE : 0));\n>      write_or_die(1, &header, ZIP_LOCAL_HEADER_SIZE);\n>      zip_offset += ZIP_LOCAL_HEADER_SIZE;\n>      write_or_die(1, path, pathlen);\n>      zip_offset += pathlen;\n>      write_or_die(1, &extra, ZIP_EXTRA_MTIME_SIZE);\n>      zip_offset += ZIP_EXTRA_MTIME_SIZE;\n> +    if (clamped) {\n> +        struct zip_extra_zip64 extra_zip64;\n> +        copy_le16(extra_zip64.magic, 0x0001);\n> +        copy_le16(extra_zip64.extra_size, ZIP_EXTRA_ZIP64_PAYLOAD_SIZE);\n> +        copy_le64(extra_zip64.size, size);\n> +        copy_le64(extra_zip64.compressed_size, compressed_size);\n> +        copy_le64(extra_zip64.offset, zip_offset);\n> +        copy_le32(extra_zip64.disk, 0);\n> +        write_or_die(1, &extra_zip64, ZIP_EXTRA_ZIP64_SIZE);\n> +        zip_offset += ZIP_EXTRA_ZIP64_SIZE;\n\nIs this correct? Not all of the zip64 extra fields are always populated. \nOnly those whose regular fields are filled with 0xffffffff must be \npresent. Since there is only one flag, it is not possible to know which \nof the fields must be filled in.\n\nReaders will most likely ignore trailing fields that should not be \nthere; however, when the offset exceeds 32 bits, but not the compressed \nsize, readers will pick the compressed size and interpret it as offset.\n\n> +    }\n\n-- Hannes\n\n"},{"id":"317551","messageId":"alpine.DEB.2.11.1704222341300.22361@perkele.intern.softwolves.pp.se","threadId":"45770","inReplyTo":"37eb7c14-eb61-7a63-bdf0-ee1ccf40723f@kdbg.org","subject":"[PATCH v2] archive-zip: Add zip64 headers when file size is too large for 32 bits","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2017-04-22T22:41:53Z","receivedAt":"2017-04-22T22:42:05Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"If the size of the files in the archive cannot be expressed in 32 bits, or\nthe offset in the zip file itself, add zip64 local headers with the actual\nsize. If we do find such entries, we also set a flag to force the creation\nof a zip64 end of central directory record.\n\nSigned-off-by: Peter Krefting <peter@softwolves.pp.se>\n---\n archive-zip.c | 50 +++++++++++++++++++++++++++++++++++++++++---------\n 1 file changed, 41 insertions(+), 9 deletions(-)\n\n> Is this correct? Not all of the zip64 extra fields are always populated. \n> Only those whose regular fields are filled with 0xffffffff must be \n> present.\n\nIndeed. Last time I implemented zip64 support it was on the reading side,\nand I remember this was a mess...\n\ndiff --git a/archive-zip.c b/archive-zip.c\nindex b429a8d97..50c7ab005 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -10,6 +10,7 @@\n \n static int zip_date;\n static int zip_time;\n+static int zip_zip64;\n \n static unsigned char *zip_dir;\n static unsigned int zip_dir_size;\n@@ -88,6 +89,16 @@ struct zip_extra_mtime {\n \tunsigned char _end[1];\n };\n \n+struct zip_extra_zip64 {\n+\tunsigned char magic[2];\n+\tunsigned char extra_size[2];\n+\tunsigned char size[8];\n+\tunsigned char compressed_size[8];\n+\tunsigned char offset[8];\n+\tunsigned char disk[4];\n+\tunsigned char _end[1];\n+};\n+\n struct zip64_dir_trailer {\n \tunsigned char magic[4];\n \tunsigned char record_size[8];\n@@ -122,6 +133,9 @@ struct zip64_dir_trailer_locator {\n #define ZIP_EXTRA_MTIME_SIZE\toffsetof(struct zip_extra_mtime, _end)\n #define ZIP_EXTRA_MTIME_PAYLOAD_SIZE \\\n \t(ZIP_EXTRA_MTIME_SIZE - offsetof(struct zip_extra_mtime, flags))\n+#define ZIP_EXTRA_ZIP64_SIZE\toffsetof(struct zip_extra_zip64, _end)\n+#define ZIP_EXTRA_ZIP64_PAYLOAD_SIZE \\\n+\t(ZIP_EXTRA_ZIP64_SIZE - offsetof(struct zip_extra_zip64, size))\n #define ZIP64_DIR_TRAILER_SIZE\toffsetof(struct zip64_dir_trailer, _end)\n #define ZIP64_DIR_TRAILER_RECORD_SIZE \\\n \t(ZIP64_DIR_TRAILER_SIZE - \\\n@@ -219,19 +233,25 @@ static void set_zip_dir_data_desc(struct zip_dir_header *header,\n \t\t\t\t  unsigned long compressed_size,\n \t\t\t\t  unsigned long crc)\n {\n+\tint clamped = 0;\n \tcopy_le32(header->crc32, crc);\n-\tcopy_le32(header->compressed_size, compressed_size);\n-\tcopy_le32(header->size, size);\n+\tcopy_le32(header->compressed_size, clamp_max(compressed_size, 0xFFFFFFFFU, &clamped));\n+\tcopy_le32(header->size, clamp_max(size, 0xFFFFFFFFU, &clamped));\n+\tif (clamped)\n+\t\tzip_zip64 = 1;\n }\n \n static void set_zip_header_data_desc(struct zip_local_header *header,\n \t\t\t\t     unsigned long size,\n \t\t\t\t     unsigned long compressed_size,\n-\t\t\t\t     unsigned long crc)\n+\t\t\t\t     unsigned long crc,\n+\t\t\t\t     int *clamped)\n {\n \tcopy_le32(header->crc32, crc);\n-\tcopy_le32(header->compressed_size, compressed_size);\n-\tcopy_le32(header->size, size);\n+\tcopy_le32(header->compressed_size, clamp_max(compressed_size, 0xFFFFFFFFU, clamped));\n+\tcopy_le32(header->size, clamp_max(size, 0xFFFFFFFFU, clamped));\n+\tif (clamped)\n+\t\tzip_zip64 = 1;\n }\n \n static int has_only_ascii(const char *s)\n@@ -279,6 +299,7 @@ static int write_zip_entry(struct archiver_args *args,\n \tint is_binary = -1;\n \tconst char *path_without_prefix = path + args->baselen;\n \tunsigned int creator_version = 0;\n+\tint clamped = 0;\n \n \tcrc = crc32(0, NULL, 0);\n \n@@ -376,7 +397,7 @@ static int write_zip_entry(struct archiver_args *args,\n \tcopy_le16(dirent.comment_length, 0);\n \tcopy_le16(dirent.disk, 0);\n \tcopy_le32(dirent.attr2, attr2);\n-\tcopy_le32(dirent.offset, zip_offset);\n+\tcopy_le32(dirent.offset, clamp_max(zip_offset, 0xFFFFFFFFU, &clamped));\n \n \tcopy_le32(header.magic, 0x04034b50);\n \tcopy_le16(header.version, 10);\n@@ -384,15 +405,26 @@ static int write_zip_entry(struct archiver_args *args,\n \tcopy_le16(header.compression_method, method);\n \tcopy_le16(header.mtime, zip_time);\n \tcopy_le16(header.mdate, zip_date);\n-\tset_zip_header_data_desc(&header, size, compressed_size, crc);\n+\tset_zip_header_data_desc(&header, size, compressed_size, crc, &clamped);\n \tcopy_le16(header.filename_length, pathlen);\n-\tcopy_le16(header.extra_length, ZIP_EXTRA_MTIME_SIZE);\n+\tcopy_le16(header.extra_length, ZIP_EXTRA_MTIME_SIZE + (clamped ? ZIP_EXTRA_ZIP64_SIZE : 0));\n \twrite_or_die(1, &header, ZIP_LOCAL_HEADER_SIZE);\n \tzip_offset += ZIP_LOCAL_HEADER_SIZE;\n \twrite_or_die(1, path, pathlen);\n \tzip_offset += pathlen;\n \twrite_or_die(1, &extra, ZIP_EXTRA_MTIME_SIZE);\n \tzip_offset += ZIP_EXTRA_MTIME_SIZE;\n+\tif (clamped) {\n+\t\tstruct zip_extra_zip64 extra_zip64;\n+\t\tcopy_le16(extra_zip64.magic, 0x0001);\n+\t\tcopy_le16(extra_zip64.extra_size, ZIP_EXTRA_ZIP64_PAYLOAD_SIZE);\n+\t\tcopy_le64(extra_zip64.size, size >= 0xFFFFFFFFU ? size : 0);\n+\t\tcopy_le64(extra_zip64.compressed_size, compressed_size >= 0xFFFFFFFFU ? compressed_size : 0);\n+\t\tcopy_le64(extra_zip64.offset, zip_offset >= 0xFFFFFFFFU ? zip_offset : 0);\n+\t\tcopy_le32(extra_zip64.disk, 0);\n+\t\twrite_or_die(1, &extra_zip64, ZIP_EXTRA_ZIP64_SIZE);\n+\t\tzip_offset += ZIP_EXTRA_ZIP64_SIZE;\n+\t}\n \tif (stream && method == 0) {\n \t\tunsigned char buf[STREAM_BUFFER_SIZE];\n \t\tssize_t readlen;\n@@ -538,7 +570,7 @@ static void write_zip_trailer(const unsigned char *sha1)\n \tcopy_le16(trailer.comment_length, sha1 ? GIT_SHA1_HEXSZ : 0);\n \n \twrite_or_die(1, zip_dir, zip_dir_offset);\n-\tif (clamped)\n+\tif (clamped || zip_zip64)\n \t\twrite_zip64_trailer();\n \twrite_or_die(1, &trailer, ZIP_DIR_TRAILER_SIZE);\n \tif (sha1)\n-- \n2.12.2\n"},{"id":"317552","messageId":"04ad7a06-969d-ffa5-b792-ccc1e7e45fd2@web.de","threadId":"45770","inReplyTo":"37eb7c14-eb61-7a63-bdf0-ee1ccf40723f@kdbg.org","subject":"Re: [PATCH] archive-zip: Add zip64 headers when file size is too large for 32 bits","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-23T00:16:59Z","receivedAt":"2017-04-23T00:17:12Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 22.04.2017 um 23:52 schrieb Johannes Sixt:\n> Am 22.04.2017 um 21:22 schrieb Peter Krefting:\n>> @@ -279,6 +299,7 @@ static int write_zip_entry(struct archiver_args \n>> *args,\n>>      int is_binary = -1;\n>>      const char *path_without_prefix = path + args->baselen;\n>>      unsigned int creator_version = 0;\n>> +    int clamped = 0;\n>>\n>>      crc = crc32(0, NULL, 0);\n>>\n>> @@ -376,7 +397,7 @@ static int write_zip_entry(struct archiver_args \n>> *args,\n>>      copy_le16(dirent.comment_length, 0);\n>>      copy_le16(dirent.disk, 0);\n>>      copy_le32(dirent.attr2, attr2);\n>> -    copy_le32(dirent.offset, zip_offset);\n>> +    copy_le32(dirent.offset, clamp_max(zip_offset, 0xFFFFFFFFU,\n>> &clamped));\n>>\n>>      copy_le32(header.magic, 0x04034b50);\n>>      copy_le16(header.version, 10);\n>> @@ -384,15 +405,26 @@ static int write_zip_entry(struct archiver_args\n>> *args,\n>>      copy_le16(header.compression_method, method);\n>>      copy_le16(header.mtime, zip_time);\n>>      copy_le16(header.mdate, zip_date);\n>> -    set_zip_header_data_desc(&header, size, compressed_size, crc);\n>> +    set_zip_header_data_desc(&header, size, compressed_size, crc,\n>> &clamped);\n>>      copy_le16(header.filename_length, pathlen);\n>> -    copy_le16(header.extra_length, ZIP_EXTRA_MTIME_SIZE);\n>> +    copy_le16(header.extra_length, ZIP_EXTRA_MTIME_SIZE + (clamped ?\n>> ZIP_EXTRA_ZIP64_SIZE : 0));\n>>      write_or_die(1, &header, ZIP_LOCAL_HEADER_SIZE);\n>>      zip_offset += ZIP_LOCAL_HEADER_SIZE;\n>>      write_or_die(1, path, pathlen);\n>>      zip_offset += pathlen;\n>>      write_or_die(1, &extra, ZIP_EXTRA_MTIME_SIZE);\n>>      zip_offset += ZIP_EXTRA_MTIME_SIZE;\n>> +    if (clamped) {\n>> +        struct zip_extra_zip64 extra_zip64;\n>> +        copy_le16(extra_zip64.magic, 0x0001);\n>> +        copy_le16(extra_zip64.extra_size, ZIP_EXTRA_ZIP64_PAYLOAD_SIZE);\n>> +        copy_le64(extra_zip64.size, size);\n>> +        copy_le64(extra_zip64.compressed_size, compressed_size);\n>> +        copy_le64(extra_zip64.offset, zip_offset);\n>> +        copy_le32(extra_zip64.disk, 0);\n>> +        write_or_die(1, &extra_zip64, ZIP_EXTRA_ZIP64_SIZE);\n>> +        zip_offset += ZIP_EXTRA_ZIP64_SIZE;\n> \n> Is this correct? Not all of the zip64 extra fields are always populated. \n> Only those whose regular fields are filled with 0xffffffff must be \n> present. Since there is only one flag, it is not possible to know which \n> of the fields must be filled in.\n> \n> Readers will most likely ignore trailing fields that should not be \n> there; however, when the offset exceeds 32 bits, but not the compressed \n> size, readers will pick the compressed size and interpret it as offset.\n\nThe offset is declared as unsigned int, so will wrap on most platforms\nbefore reaching the clamp check.  At least InfoZIP's unzip can handle\nthat, but it's untidy.\n\nThe offset is only needed in the ZIP64 extra record for the central\nheader (in zip_dir) -- the local header has no offset field.  That said,\nI haven't been able to implement proper 64 bit offset support so far.\n\nRené\n"},{"id":"317556","messageId":"alpine.DEB.2.11.1704230737380.29888@perkele.intern.softwolves.pp.se","threadId":"45770","inReplyTo":"04ad7a06-969d-ffa5-b792-ccc1e7e45fd2@web.de","subject":"Re: [PATCH] archive-zip: Add zip64 headers when file size is too large for 32 bits","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2017-04-23T06:42:37Z","receivedAt":"2017-04-23T06:43:23Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"René Scharfe:\n\n> The offset is declared as unsigned int, so will wrap on most platforms\n> before reaching the clamp check.  At least InfoZIP's unzip can handle\n> that, but it's untidy.\n\nRight, that needs to be changed into unsigned long and clamped, just \nlike the original and compressed file sizes already are.\n\n> The offset is only needed in the ZIP64 extra record for the central \n> header (in zip_dir) -- the local header has no offset field.\n\nThe zip64 local header does have an offset field, though. I thought \nthat was the zip_offset value, but that doesn't make sense, I'm not \nquite sure what it is supposed to store. I need to investigate that \nfurther, I assume.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"317557","messageId":"daa66f7e-b77e-3a27-a6f1-7d9059ab71b7@kdbg.org","threadId":"45770","inReplyTo":"alpine.DEB.2.11.1704230737380.29888@perkele.intern.softwolves.pp.se","subject":"Re: [PATCH] archive-zip: Add zip64 headers when file size is too large for 32 bits","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2017-04-23T07:27:23Z","receivedAt":"2017-04-23T07:27:31Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 23.04.2017 um 08:42 schrieb Peter Krefting:\n> René Scharfe:\n>> The offset is only needed in the ZIP64 extra record for the central\n>> header (in zip_dir) -- the local header has no offset field.\n\nGood point.\n\n> The zip64 local header does have an offset field, though. I thought that\n> was the zip_offset value, but that doesn't make sense, I'm not quite\n> sure what it is supposed to store. I need to investigate that further, I\n> assume.\n\nLet's get the naming straight: There is no \"zip64 local header\". There \nis a \"zip64 extra record\" for the \"zip local header\". The zip64 extra \ndata record has an offset field, but since the local header does not \nhave an offset field, the offset field in the corresponding zip64 extra \ndata record is always omitted.\n\n-- Hannes\n\n"},{"id":"317558","messageId":"a1504d15-36d6-51f8-f2c9-a6563789bb6f@kdbg.org","threadId":"45770","inReplyTo":"alpine.DEB.2.11.1704222341300.22361@perkele.intern.softwolves.pp.se","subject":"Re: [PATCH v2] archive-zip: Add zip64 headers when file size is too large for 32 bits","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2017-04-23T07:50:10Z","receivedAt":"2017-04-23T07:50:24Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 23.04.2017 um 00:41 schrieb Peter Krefting:\n> Indeed. Last time I implemented zip64 support it was on the reading side,\n> and I remember this was a mess...\n\nIt is indeed!\n\n>  static void set_zip_header_data_desc(struct zip_local_header *header,\n>  \t\t\t\t     unsigned long size,\n>  \t\t\t\t     unsigned long compressed_size,\n> -\t\t\t\t     unsigned long crc)\n> +\t\t\t\t     unsigned long crc,\n> +\t\t\t\t     int *clamped)\n>  {\n>  \tcopy_le32(header->crc32, crc);\n> -\tcopy_le32(header->compressed_size, compressed_size);\n> -\tcopy_le32(header->size, size);\n> +\tcopy_le32(header->compressed_size, clamp_max(compressed_size, 0xFFFFFFFFU, clamped));\n> +\tcopy_le32(header->size, clamp_max(size, 0xFFFFFFFFU, clamped));\n> +\tif (clamped)\n> +\t\tzip_zip64 = 1;\n\nThis must be\n\n\tif (*clamped)\n\n>  }\n>\n>  static int has_only_ascii(const char *s)\n> @@ -279,6 +299,7 @@ static int write_zip_entry(struct archiver_args *args,\n>  \tint is_binary = -1;\n>  \tconst char *path_without_prefix = path + args->baselen;\n>  \tunsigned int creator_version = 0;\n> +\tint clamped = 0;\n>\n>  \tcrc = crc32(0, NULL, 0);\n>\n> @@ -376,7 +397,7 @@ static int write_zip_entry(struct archiver_args *args,\n>  \tcopy_le16(dirent.comment_length, 0);\n>  \tcopy_le16(dirent.disk, 0);\n>  \tcopy_le32(dirent.attr2, attr2);\n> -\tcopy_le32(dirent.offset, zip_offset);\n> +\tcopy_le32(dirent.offset, clamp_max(zip_offset, 0xFFFFFFFFU, &clamped));\n\nI don't see any provisions to write the zip64 extra header in the \ncentral directory when this offset is clamped. This means that ZIP \narchives whose size exceed 4GB are still unsupported.\n\n>\n>  \tcopy_le32(header.magic, 0x04034b50);\n>  \tcopy_le16(header.version, 10);\n> @@ -384,15 +405,26 @@ static int write_zip_entry(struct archiver_args *args,\n>  \tcopy_le16(header.compression_method, method);\n>  \tcopy_le16(header.mtime, zip_time);\n>  \tcopy_le16(header.mdate, zip_date);\n> -\tset_zip_header_data_desc(&header, size, compressed_size, crc);\n> +\tset_zip_header_data_desc(&header, size, compressed_size, crc, &clamped);\n>  \tcopy_le16(header.filename_length, pathlen);\n> -\tcopy_le16(header.extra_length, ZIP_EXTRA_MTIME_SIZE);\n> +\tcopy_le16(header.extra_length, ZIP_EXTRA_MTIME_SIZE + (clamped ? ZIP_EXTRA_ZIP64_SIZE : 0));\n>  \twrite_or_die(1, &header, ZIP_LOCAL_HEADER_SIZE);\n>  \tzip_offset += ZIP_LOCAL_HEADER_SIZE;\n>  \twrite_or_die(1, path, pathlen);\n>  \tzip_offset += pathlen;\n>  \twrite_or_die(1, &extra, ZIP_EXTRA_MTIME_SIZE);\n>  \tzip_offset += ZIP_EXTRA_MTIME_SIZE;\n> +\tif (clamped) {\n> +\t\tstruct zip_extra_zip64 extra_zip64;\n> +\t\tcopy_le16(extra_zip64.magic, 0x0001);\n> +\t\tcopy_le16(extra_zip64.extra_size, ZIP_EXTRA_ZIP64_PAYLOAD_SIZE);\n> +\t\tcopy_le64(extra_zip64.size, size >= 0xFFFFFFFFU ? size : 0);\n> +\t\tcopy_le64(extra_zip64.compressed_size, compressed_size >= 0xFFFFFFFFU ? compressed_size : 0);\n> +\t\tcopy_le64(extra_zip64.offset, zip_offset >= 0xFFFFFFFFU ? zip_offset : 0);\n> +\t\tcopy_le32(extra_zip64.disk, 0);\n\nThese are wrong, I think. Entries that did not overflow must be omitted \nentirely from the zip64 extra record, not filled with 0. This implies \nthat the payload size (.extra_size) is dynamic.\n\nAs René pointed out, the offset is only written in the central \ndirectory, but not in the local header for the current file. Therefore, \nit must be omitted here. The disk number also never exceeds 0xffff and \nmust be omitted as well.\n\n> +\t\twrite_or_die(1, &extra_zip64, ZIP_EXTRA_ZIP64_SIZE);\n> +\t\tzip_offset += ZIP_EXTRA_ZIP64_SIZE;\n> +\t}\n>  \tif (stream && method == 0) {\n>  \t\tunsigned char buf[STREAM_BUFFER_SIZE];\n>  \t\tssize_t readlen;\n> @@ -538,7 +570,7 @@ static void write_zip_trailer(const unsigned char *sha1)\n>  \tcopy_le16(trailer.comment_length, sha1 ? GIT_SHA1_HEXSZ : 0);\n>\n>  \twrite_or_die(1, zip_dir, zip_dir_offset);\n> -\tif (clamped)\n> +\tif (clamped || zip_zip64)\n>  \t\twrite_zip64_trailer();\n>  \twrite_or_die(1, &trailer, ZIP_DIR_TRAILER_SIZE);\n>  \tif (sha1)\n>\n\n-- Hannes\n\n"},{"id":"317561","messageId":"alpine.DEB.2.11.1704231526450.3944@perkele.intern.softwolves.pp.se","threadId":"45770","inReplyTo":"a1504d15-36d6-51f8-f2c9-a6563789bb6f@kdbg.org","subject":"Re: [PATCH v2] archive-zip: Add zip64 headers when file size is too large for 32 bits","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2017-04-23T14:51:21Z","receivedAt":"2017-04-23T14:51:28Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Johannes Sixt:\n\n> Let's get the naming straight: There is no \"zip64 local header\". There is a \n> \"zip64 extra record\" for the \"zip local header\".\n\nIndeed, sorry for the confusion. That's what I get for trying to write \ncoherent email at half past midnight :-)\n\n> The zip64 extra data record has an offset field, but since the local \n> header does not have an offset field, the offset field in the \n> corresponding zip64 extra data record is always omitted.\n\nAh, now I understand, I was a bit confused, as the same code generates \nthe central directory entry as the local entry.\n\n>> @@ -376,7 +397,7 @@ static int write_zip_entry(struct archiver_args *args,\n>>  \tcopy_le16(dirent.comment_length, 0);\n>>  \tcopy_le16(dirent.disk, 0);\n>>  \tcopy_le32(dirent.attr2, attr2);\n>> -\tcopy_le32(dirent.offset, zip_offset);\n>> +\tcopy_le32(dirent.offset, clamp_max(zip_offset, 0xFFFFFFFFU, \n>> &clamped));\n>\n> I don't see any provisions to write the zip64 extra header in the central \n> directory when this offset is clamped. This means that ZIP archives whose \n> size exceed 4GB are still unsupported.\n\nThe clamped flag will trigger the inclusion of the zip64 central \ndirectory, which contains the 64-bit offset. Should the central \ndirectory entry also have the zip64 extra field?\n\n> These are wrong, I think. Entries that did not overflow must be omitted \n> entirely from the zip64 extra record, not filled with 0. This implies that \n> the payload size (.extra_size) is dynamic.\n\nThat is what I was trying to figure out, APPNOTE is extremely vague on \nthe subject, but thinking back I recall that you are correct.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"317566","messageId":"e0d1c923-a9f5-9ffc-a7e7-67f558e50796@kdbg.org","threadId":"45770","inReplyTo":"alpine.DEB.2.11.1704231526450.3944@perkele.intern.softwolves.pp.se","subject":"Re: [PATCH v2] archive-zip: Add zip64 headers when file size is too large for 32 bits","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2017-04-23T19:49:53Z","receivedAt":"2017-04-23T19:50:02Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 23.04.2017 um 16:51 schrieb Peter Krefting:\n> Johannes Sixt:\n>>> @@ -376,7 +397,7 @@ static int write_zip_entry(struct archiver_args\n>>> *args,\n>>>      copy_le16(dirent.comment_length, 0);\n>>>      copy_le16(dirent.disk, 0);\n>>>      copy_le32(dirent.attr2, attr2);\n>>> -    copy_le32(dirent.offset, zip_offset);\n>>> +    copy_le32(dirent.offset, clamp_max(zip_offset, 0xFFFFFFFFU,\n>>> &clamped));\n>>\n>> I don't see any provisions to write the zip64 extra header in the\n>> central directory when this offset is clamped. This means that ZIP\n>> archives whose size exceed 4GB are still unsupported.\n>\n> The clamped flag will trigger the inclusion of the zip64 central\n> directory, which contains the 64-bit offset. Should the central\n> directory entry also have the zip64 extra field?\n\nThere is no \"zip64 central directory\". There is a \"zip64 end of central \ndirectory record\"; it tells where to find the \"central directory\" in \ncase that the ZIP archive exceeds 4GB. The central directory has the \nsame format in a non-zip64 and a zip64 archive. But when size, \ncompressed size, and offset overflow 4GB, it uses the same zip64 extra \nrecord like the local header records, except that it has to record an \noffset in addition to the uncompressed and compressed sizes.\n\nThe uncompressed and compressed sizes of entries are mentioned in both \nthe central directory and the individual local headers. I think that the \ncentral directory's values are authorative; my reasoning is that it is \npossible that the local header can have a bit set that tells that the \nlocal header's values size values are garbage.\n\nIn summary, yes, when the central directory is constructed, it must use \nthe zip64 extra record to note the values of uncompressed size, \ncompressed size, and the offset to the local header when they overflow 4GB.\n\n-- Hannes\n\n"},{"id":"317650","messageId":"alpine.DEB.2.00.1704240901520.31537@ds9.cixit.se","threadId":"45770","inReplyTo":"e0d1c923-a9f5-9ffc-a7e7-67f558e50796@kdbg.org","subject":"Re: [PATCH v2] archive-zip: Add zip64 headers when file size is too large for 32 bits","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2017-04-24T08:04:08Z","receivedAt":"2017-04-24T08:11:53Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Johannes Sixt:\n\n> There is no \"zip64 central directory\". There is a \"zip64 end of central \n> directory record\";\n\nNot strange that I was confused and couldn't find it, then... :-)\n\nAll right, I need to fix up my patch to make sure I add the zip64 \nextra record to both the central directory entry and to the local \nheader, and make sure to trigger the zip64 end of central directory \nwhenever the zip file is large enough to warrant one.\n\n> In summary, yes, when the central directory is constructed, it must \n> use the zip64 extra record to note the values of uncompressed size, \n> compressed size, and the offset to the local header when they \n> overflow 4GB.\n\nAt least that makes it easier to construct, as we only have one \ncentral directory and can just extend the records that need extending.\n\nWill fix soon.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"317673","messageId":"b3f2f12c-2736-46ed-62c9-16334c5e3483@web.de","threadId":"45770","inReplyTo":"alpine.DEB.2.00.1704240901520.31537@ds9.cixit.se","subject":"Re: [PATCH v2] archive-zip: Add zip64 headers when file size is too large for 32 bits","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-24T12:04:12Z","receivedAt":"2017-04-24T12:04:25Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 24.04.2017 um 10:04 schrieb Peter Krefting:\n> Johannes Sixt:\n> \n>> There is no \"zip64 central directory\". There is a \"zip64 end of \n>> central directory record\";\n> \n> Not strange that I was confused and couldn't find it, then... :-)\n> \n> All right, I need to fix up my patch to make sure I add the zip64 extra \n> record to both the central directory entry and to the local header, and \n> make sure to trigger the zip64 end of central directory whenever the zip \n> file is large enough to warrant one.\n> \n>> In summary, yes, when the central directory is constructed, it must \n>> use the zip64 extra record to note the values of uncompressed size, \n>> compressed size, and the offset to the local header when they overflow \n>> 4GB.\n> \n> At least that makes it easier to construct, as we only have one central \n> directory and can just extend the records that need extending.\n> \n> Will fix soon.\n\nI have a few patches for that as well.  Testing in particular is a bit \ntricky -- how to avoid creating multi-GB files just for this small \nfeature?  I hope to have something to show later today.\n\nRené\n"},{"id":"317702","messageId":"85f2b6d1-107b-0624-af82-92446f28269e@web.de","threadId":"45770","inReplyTo":"b3f2f12c-2736-46ed-62c9-16334c5e3483@web.de","subject":"[PATCH v3 0/5] archive-zip: support files and archives bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-24T17:22:03Z","receivedAt":"2017-04-24T17:22:15Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"The first patch adds (expensive) tests, the next two are cleanups which\nset the stage for the remaining two to actually implement zip64 support\nfor offsets and file sizes.\n\nHalf of the series had been laying around for months, half-finished and\nforgotten because I got distracted by the holiday season. :-/\n\n  archive-zip: add tests for big ZIP archives\n  archive-zip: use strbuf for ZIP directory\n  archive-zip: write ZIP dir entry directly to strbuf\n  archive-zip: support archives bigger than 4GB\n  archive-zip: support files bigger than 4GB\n\n archive-zip.c                   | 211 ++++++++++++++++++++++++----------------\n t/t5004-archive-corner-cases.sh |  45 +++++++++\n t/t5004/big-pack.zip            | Bin 0 -> 7373 bytes\n 3 files changed, 172 insertions(+), 84 deletions(-)\n create mode 100644 t/t5004/big-pack.zip\n\n-- \n2.12.2\n\n"},{"id":"317704","messageId":"7361b84d-41c4-8224-ff83-36703837eb28@web.de","threadId":"45770","inReplyTo":"85f2b6d1-107b-0624-af82-92446f28269e@web.de","subject":"[PATCH v3 1/5] archive-zip: add tests for big ZIP archives","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-24T17:29:53Z","receivedAt":"2017-04-24T17:30:06Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Test the creation of ZIP archives bigger than 4GB and containing files\nbigger than 4GB.  They are marked as EXPENSIVE because they take quite a\nwhile and because the first one needs a bit more than 4GB of disk space\nto store the resulting archive.\n\nThe big archive in the first test is made up of a tree containing\nthousands of copies of a small file.  Yet the test has to write out the\nfull archive because unzip doesn't offer a way to read from stdin.\n\nThe big file in the second test is provided as a zipped pack file to\navoid writing another 4GB file to disk and then adding it.\n\nSigned-off-by: Rene Scharfe <l.s.r@web.de>\n---\n t/t5004-archive-corner-cases.sh |  45 ++++++++++++++++++++++++++++++++++++++++\n t/t5004/big-pack.zip            | Bin 0 -> 7373 bytes\n 2 files changed, 45 insertions(+)\n create mode 100644 t/t5004/big-pack.zip\n\ndiff --git a/t/t5004-archive-corner-cases.sh b/t/t5004-archive-corner-cases.sh\nindex cca23383c5..bc052c803a 100755\n--- a/t/t5004-archive-corner-cases.sh\n+++ b/t/t5004-archive-corner-cases.sh\n@@ -155,4 +155,49 @@ test_expect_success ZIPINFO 'zip archive with many entries' '\n \ttest_cmp expect actual\n '\n \n+test_expect_failure EXPENSIVE,UNZIP 'zip archive bigger than 4GB' '\n+\t# build string containing 65536 characters\n+\ts=0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef &&\n+\ts=$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s &&\n+\ts=$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s &&\n+\n+\t# create blob with a length of 65536 + 1 bytes\n+\tblob=$(echo $s | git hash-object -w --stdin) &&\n+\n+\t# create tree containing 65500 entries of that blob\n+\tfor i in $(test_seq 1 65500)\n+\tdo\n+\t\techo \"100644 blob $blob\t$i\"\n+\tdone >tree &&\n+\ttree=$(git mktree <tree) &&\n+\n+\t# zip it, creating an archive a bit bigger than 4GB\n+\tgit archive -0 -o many-big.zip $tree &&\n+\n+\t\"$GIT_UNZIP\" -t many-big.zip 9999 65500 &&\n+\t\"$GIT_UNZIP\" -t many-big.zip\n+'\n+\n+test_expect_failure EXPENSIVE,UNZIP,ZIPINFO 'zip archive with files bigger than 4GB' '\n+\t# Pack created with:\n+\t#   dd if=/dev/zero of=file bs=1M count=4100 && git hash-object -w file\n+\tmkdir -p .git/objects/pack &&\n+\t(\n+\t\tcd .git/objects/pack &&\n+\t\t\"$GIT_UNZIP\" \"$TEST_DIRECTORY\"/t5004/big-pack.zip\n+\t) &&\n+\tblob=754a93d6fada4c6873360e6cb4b209132271ab0e &&\n+\tsize=$(expr 4100 \"*\" 1024 \"*\" 1024) &&\n+\n+\t# create a tree containing the file\n+\ttree=$(echo \"100644 blob $blob\tbig-file\" | git mktree) &&\n+\n+\t# zip it, creating an archive with a file bigger than 4GB\n+\tgit archive -o big.zip $tree &&\n+\n+\t\"$GIT_UNZIP\" -t big.zip &&\n+\t\"$ZIPINFO\" big.zip >big.lst &&\n+\tgrep $size big.lst\n+'\n+\n test_done\ndiff --git a/t/t5004/big-pack.zip b/t/t5004/big-pack.zip\nnew file mode 100644\nindex 0000000000000000000000000000000000000000..caaf614eeece6f818c525e433561e37560a75b05\nGIT binary patch\nliteral 7373\nzcmWIWW@Zs#U}E54xLY#A%SmVLl4u471|Jp%215oJhJwW8Y+ZvC^HlQ`3(KT5gS6zd\nz#6$y=#3VyYlN6KGlw?bTG()50MB`LTb5p&{l#0+0P6p<OAOA*xaA^fM10%}|W(Ec@\nz@jL#}{4)m*95`aY)ux;vb9C_xiI`WWKX0v%wv#-XR%01$R>-9-y6#!_6|EcGQyxED\nzyZY^&^YwM2r?Wic_gf3wz0#?R{8hB^<Js4()@9Ms>$gt{iHiUEJN{iALjZ~|R(VzA\nzOp{_@IDE*S!H8sEfc%Wl8*lHd?-DJPX@5BLD~n)>se}uQ{sHa{9CBhhi*hb3*-*jg\nzc5o4&4%_XW4Ba<OG7P)VrhVaOTdmN<Ho4-C?R3NUpAYE!{5{5hc=zr*JFR#QUppur\nz?|(T<>b?Enk9jiu`<7qbA8y|-Ck>2+hMBYN-pgCcFt>ev-5*|-&jb`JE>Hh`B*~cT\nz&t2=?C44}E8GDaU*PD0ekK%{l7w@(f14RzJnfvea+fQlS75nRoU$s?(g=&R(fOLb%\nzK_JPHAqeJ(fjJ$>oKcz4&>2l3kc=^!7e@2KXkHl23k(gTVK5p7z>;7z9gKznQp0()\nzeK6WS7;PVn){Ud}!f4$%I)=h9I*!CJ8V10UU^E?!h5@KqG@1@Z!(cQWK$^#7<%OS{\nz_HQoT!K!n;di%NDON?i(ygj-((dOTWgOj)4mXFhko42#<@9S6BTc6*5|Cc?$n~_P5\nz8P`mn1SleafRRC^5k!+Qug40R*F&4rL$?-n>J8c2V<cM(nTW$>FDo0!BTPW}9!Q@6\nI&6hC%0B76lcmMzZ\n\nliteral 0\nHcmV?d00001\n\n-- \n2.12.2\n\n"},{"id":"317705","messageId":"949f19e6-0414-9abc-9754-064d7e58c169@web.de","threadId":"45770","inReplyTo":"85f2b6d1-107b-0624-af82-92446f28269e@web.de","subject":"[PATCH v3 2/5] archive-zip: use strbuf for ZIP directory","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-24T17:30:49Z","receivedAt":"2017-04-24T17:30:59Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Keep the ZIP central director, which is written after all archive\nentries, in a strbuf instead of a custom-managed buffer.  It contains\nbinary data, so we can't (and don't want to) use the full range of\nstrbuf functions and we don't need the terminating NUL, but the result\nis shorter and simpler code.\n\nSigned-off-by: Rene Scharfe <l.s.r@web.de>\n---\n archive-zip.c | 36 +++++++++++-------------------------\n 1 file changed, 11 insertions(+), 25 deletions(-)\n\ndiff --git a/archive-zip.c b/archive-zip.c\nindex b429a8d974..a6fac59602 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -11,16 +11,14 @@\n static int zip_date;\n static int zip_time;\n \n-static unsigned char *zip_dir;\n-static unsigned int zip_dir_size;\n+/* We only care about the \"buf\" part here. */\n+static struct strbuf zip_dir;\n \n static unsigned int zip_offset;\n-static unsigned int zip_dir_offset;\n static uint64_t zip_dir_entries;\n \n static unsigned int max_creator_version;\n \n-#define ZIP_DIRECTORY_MIN_SIZE\t(1024 * 1024)\n #define ZIP_STREAM\t(1 <<  3)\n #define ZIP_UTF8\t(1 << 11)\n \n@@ -268,7 +266,6 @@ static int write_zip_entry(struct archiver_args *args,\n \tunsigned long attr2;\n \tunsigned long compressed_size;\n \tunsigned long crc;\n-\tunsigned long direntsize;\n \tint method;\n \tunsigned char *out;\n \tvoid *deflated = NULL;\n@@ -356,13 +353,6 @@ static int write_zip_entry(struct archiver_args *args,\n \textra.flags[0] = 1;\t/* just mtime */\n \tcopy_le32(extra.mtime, args->time);\n \n-\t/* make sure we have enough free space in the dictionary */\n-\tdirentsize = ZIP_DIR_HEADER_SIZE + pathlen + ZIP_EXTRA_MTIME_SIZE;\n-\twhile (zip_dir_size < zip_dir_offset + direntsize) {\n-\t\tzip_dir_size += ZIP_DIRECTORY_MIN_SIZE;\n-\t\tzip_dir = xrealloc(zip_dir, zip_dir_size);\n-\t}\n-\n \tcopy_le32(dirent.magic, 0x02014b50);\n \tcopy_le16(dirent.creator_version, creator_version);\n \tcopy_le16(dirent.version, 10);\n@@ -486,12 +476,9 @@ static int write_zip_entry(struct archiver_args *args,\n \n \tcopy_le16(dirent.attr1, !is_binary);\n \n-\tmemcpy(zip_dir + zip_dir_offset, &dirent, ZIP_DIR_HEADER_SIZE);\n-\tzip_dir_offset += ZIP_DIR_HEADER_SIZE;\n-\tmemcpy(zip_dir + zip_dir_offset, path, pathlen);\n-\tzip_dir_offset += pathlen;\n-\tmemcpy(zip_dir + zip_dir_offset, &extra, ZIP_EXTRA_MTIME_SIZE);\n-\tzip_dir_offset += ZIP_EXTRA_MTIME_SIZE;\n+\tstrbuf_add(&zip_dir, &dirent, ZIP_DIR_HEADER_SIZE);\n+\tstrbuf_add(&zip_dir, path, pathlen);\n+\tstrbuf_add(&zip_dir, &extra, ZIP_EXTRA_MTIME_SIZE);\n \tzip_dir_entries++;\n \n \treturn 0;\n@@ -510,12 +497,12 @@ static void write_zip64_trailer(void)\n \tcopy_le32(trailer64.directory_start_disk, 0);\n \tcopy_le64(trailer64.entries_on_this_disk, zip_dir_entries);\n \tcopy_le64(trailer64.entries, zip_dir_entries);\n-\tcopy_le64(trailer64.size, zip_dir_offset);\n+\tcopy_le64(trailer64.size, zip_dir.len);\n \tcopy_le64(trailer64.offset, zip_offset);\n \n \tcopy_le32(locator64.magic, 0x07064b50);\n \tcopy_le32(locator64.disk, 0);\n-\tcopy_le64(locator64.offset, zip_offset + zip_dir_offset);\n+\tcopy_le64(locator64.offset, zip_offset + zip_dir.len);\n \tcopy_le32(locator64.number_of_disks, 1);\n \n \twrite_or_die(1, &trailer64, ZIP64_DIR_TRAILER_SIZE);\n@@ -533,11 +520,11 @@ static void write_zip_trailer(const unsigned char *sha1)\n \tcopy_le16_clamp(trailer.entries_on_this_disk, zip_dir_entries,\n \t\t\t&clamped);\n \tcopy_le16_clamp(trailer.entries, zip_dir_entries, &clamped);\n-\tcopy_le32(trailer.size, zip_dir_offset);\n+\tcopy_le32(trailer.size, zip_dir.len);\n \tcopy_le32(trailer.offset, zip_offset);\n \tcopy_le16(trailer.comment_length, sha1 ? GIT_SHA1_HEXSZ : 0);\n \n-\twrite_or_die(1, zip_dir, zip_dir_offset);\n+\twrite_or_die(1, zip_dir.buf, zip_dir.len);\n \tif (clamped)\n \t\twrite_zip64_trailer();\n \twrite_or_die(1, &trailer, ZIP_DIR_TRAILER_SIZE);\n@@ -568,14 +555,13 @@ static int write_zip_archive(const struct archiver *ar,\n \n \tdos_time(&args->time, &zip_date, &zip_time);\n \n-\tzip_dir = xmalloc(ZIP_DIRECTORY_MIN_SIZE);\n-\tzip_dir_size = ZIP_DIRECTORY_MIN_SIZE;\n+\tstrbuf_init(&zip_dir, 0);\n \n \terr = write_archive_entries(args, write_zip_entry);\n \tif (!err)\n \t\twrite_zip_trailer(args->commit_sha1);\n \n-\tfree(zip_dir);\n+\tstrbuf_release(&zip_dir);\n \n \treturn err;\n }\n-- \n2.12.2\n\n"},{"id":"317706","messageId":"dda8fd49-87af-057d-d929-3240f8fb3d37@web.de","threadId":"45770","inReplyTo":"85f2b6d1-107b-0624-af82-92446f28269e@web.de","subject":"[PATCH v3 3/5] archive-zip: write ZIP dir entry directly to strbuf","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-24T17:31:44Z","receivedAt":"2017-04-24T17:31:52Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Write all fields of the ZIP directory record for an archive entry\nin the right order directly into the strbuf instead of taking a detour\nthrough a struct.  Do that at end, when we have all necessary data like\nchecksum and compressed size.  The fields are documented just as well,\nthe code becomes shorter and we save an extra copy.\n\nSigned-off-by: Rene Scharfe <l.s.r@web.de>\n---\n archive-zip.c | 81 ++++++++++++++++++++---------------------------------------\n 1 file changed, 27 insertions(+), 54 deletions(-)\n\ndiff --git a/archive-zip.c b/archive-zip.c\nindex a6fac59602..2d52bb3ade 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -45,27 +45,6 @@ struct zip_data_desc {\n \tunsigned char _end[1];\n };\n \n-struct zip_dir_header {\n-\tunsigned char magic[4];\n-\tunsigned char creator_version[2];\n-\tunsigned char version[2];\n-\tunsigned char flags[2];\n-\tunsigned char compression_method[2];\n-\tunsigned char mtime[2];\n-\tunsigned char mdate[2];\n-\tunsigned char crc32[4];\n-\tunsigned char compressed_size[4];\n-\tunsigned char size[4];\n-\tunsigned char filename_length[2];\n-\tunsigned char extra_length[2];\n-\tunsigned char comment_length[2];\n-\tunsigned char disk[2];\n-\tunsigned char attr1[2];\n-\tunsigned char attr2[4];\n-\tunsigned char offset[4];\n-\tunsigned char _end[1];\n-};\n-\n struct zip_dir_trailer {\n \tunsigned char magic[4];\n \tunsigned char disk[2];\n@@ -166,6 +145,15 @@ static void copy_le16_clamp(unsigned char *dest, uint64_t n, int *clamped)\n \tcopy_le16(dest, clamp_max(n, 0xffff, clamped));\n }\n \n+static int strbuf_add_le(struct strbuf *sb, size_t size, uintmax_t n)\n+{\n+\twhile (size-- > 0) {\n+\t\tstrbuf_addch(sb, n & 0xff);\n+\t\tn >>= 8;\n+\t}\n+\treturn -!!n;\n+}\n+\n static void *zlib_deflate_raw(void *data, unsigned long size,\n \t\t\t      int compression_level,\n \t\t\t      unsigned long *compressed_size)\n@@ -212,16 +200,6 @@ static void write_zip_data_desc(unsigned long size,\n \twrite_or_die(1, &trailer, ZIP_DATA_DESC_SIZE);\n }\n \n-static void set_zip_dir_data_desc(struct zip_dir_header *header,\n-\t\t\t\t  unsigned long size,\n-\t\t\t\t  unsigned long compressed_size,\n-\t\t\t\t  unsigned long crc)\n-{\n-\tcopy_le32(header->crc32, crc);\n-\tcopy_le32(header->compressed_size, compressed_size);\n-\tcopy_le32(header->size, size);\n-}\n-\n static void set_zip_header_data_desc(struct zip_local_header *header,\n \t\t\t\t     unsigned long size,\n \t\t\t\t     unsigned long compressed_size,\n@@ -261,7 +239,7 @@ static int write_zip_entry(struct archiver_args *args,\n \t\t\t   unsigned int mode)\n {\n \tstruct zip_local_header header;\n-\tstruct zip_dir_header dirent;\n+\tuintmax_t offset = zip_offset;\n \tstruct zip_extra_mtime extra;\n \tunsigned long attr2;\n \tunsigned long compressed_size;\n@@ -353,21 +331,6 @@ static int write_zip_entry(struct archiver_args *args,\n \textra.flags[0] = 1;\t/* just mtime */\n \tcopy_le32(extra.mtime, args->time);\n \n-\tcopy_le32(dirent.magic, 0x02014b50);\n-\tcopy_le16(dirent.creator_version, creator_version);\n-\tcopy_le16(dirent.version, 10);\n-\tcopy_le16(dirent.flags, flags);\n-\tcopy_le16(dirent.compression_method, method);\n-\tcopy_le16(dirent.mtime, zip_time);\n-\tcopy_le16(dirent.mdate, zip_date);\n-\tset_zip_dir_data_desc(&dirent, size, compressed_size, crc);\n-\tcopy_le16(dirent.filename_length, pathlen);\n-\tcopy_le16(dirent.extra_length, ZIP_EXTRA_MTIME_SIZE);\n-\tcopy_le16(dirent.comment_length, 0);\n-\tcopy_le16(dirent.disk, 0);\n-\tcopy_le32(dirent.attr2, attr2);\n-\tcopy_le32(dirent.offset, zip_offset);\n-\n \tcopy_le32(header.magic, 0x04034b50);\n \tcopy_le16(header.version, 10);\n \tcopy_le16(header.flags, flags);\n@@ -406,8 +369,6 @@ static int write_zip_entry(struct archiver_args *args,\n \n \t\twrite_zip_data_desc(size, compressed_size, crc);\n \t\tzip_offset += ZIP_DATA_DESC_SIZE;\n-\n-\t\tset_zip_dir_data_desc(&dirent, size, compressed_size, crc);\n \t} else if (stream && method == 8) {\n \t\tunsigned char buf[STREAM_BUFFER_SIZE];\n \t\tssize_t readlen;\n@@ -464,8 +425,6 @@ static int write_zip_entry(struct archiver_args *args,\n \n \t\twrite_zip_data_desc(size, compressed_size, crc);\n \t\tzip_offset += ZIP_DATA_DESC_SIZE;\n-\n-\t\tset_zip_dir_data_desc(&dirent, size, compressed_size, crc);\n \t} else if (compressed_size > 0) {\n \t\twrite_or_die(1, out, compressed_size);\n \t\tzip_offset += compressed_size;\n@@ -474,9 +433,23 @@ static int write_zip_entry(struct archiver_args *args,\n \tfree(deflated);\n \tfree(buffer);\n \n-\tcopy_le16(dirent.attr1, !is_binary);\n-\n-\tstrbuf_add(&zip_dir, &dirent, ZIP_DIR_HEADER_SIZE);\n+\tstrbuf_add_le(&zip_dir, 4, 0x02014b50);\t/* magic */\n+\tstrbuf_add_le(&zip_dir, 2, creator_version);\n+\tstrbuf_add_le(&zip_dir, 2, 10);\t\t/* version */\n+\tstrbuf_add_le(&zip_dir, 2, flags);\n+\tstrbuf_add_le(&zip_dir, 2, method);\n+\tstrbuf_add_le(&zip_dir, 2, zip_time);\n+\tstrbuf_add_le(&zip_dir, 2, zip_date);\n+\tstrbuf_add_le(&zip_dir, 4, crc);\n+\tstrbuf_add_le(&zip_dir, 4, compressed_size);\n+\tstrbuf_add_le(&zip_dir, 4, size);\n+\tstrbuf_add_le(&zip_dir, 2, pathlen);\n+\tstrbuf_add_le(&zip_dir, 2, ZIP_EXTRA_MTIME_SIZE);\n+\tstrbuf_add_le(&zip_dir, 2, 0);\t\t/* comment length */\n+\tstrbuf_add_le(&zip_dir, 2, 0);\t\t/* disk */\n+\tstrbuf_add_le(&zip_dir, 2, !is_binary);\n+\tstrbuf_add_le(&zip_dir, 4, attr2);\n+\tstrbuf_add_le(&zip_dir, 4, offset);\n \tstrbuf_add(&zip_dir, path, pathlen);\n \tstrbuf_add(&zip_dir, &extra, ZIP_EXTRA_MTIME_SIZE);\n \tzip_dir_entries++;\n-- \n2.12.2\n\n"},{"id":"317707","messageId":"02ddca3c-a11f-7c0c-947e-5ca87a62cdee@web.de","threadId":"45770","inReplyTo":"85f2b6d1-107b-0624-af82-92446f28269e@web.de","subject":"[PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-24T17:32:36Z","receivedAt":"2017-04-24T17:32:45Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Add a zip64 extended information extra field to the central directory\nand emit the zip64 end of central directory records as well as locator\nif the offset of an entry within the archive exceeds 4GB.\n\nSigned-off-by: Rene Scharfe <l.s.r@web.de>\n---\n archive-zip.c                   | 32 ++++++++++++++++++++++++++++----\n t/t5004-archive-corner-cases.sh |  2 +-\n 2 files changed, 29 insertions(+), 5 deletions(-)\n\ndiff --git a/archive-zip.c b/archive-zip.c\nindex 2d52bb3ade..7d6f2a85d0 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -14,7 +14,7 @@ static int zip_time;\n /* We only care about the \"buf\" part here. */\n static struct strbuf zip_dir;\n \n-static unsigned int zip_offset;\n+static uintmax_t zip_offset;\n static uint64_t zip_dir_entries;\n \n static unsigned int max_creator_version;\n@@ -145,6 +145,11 @@ static void copy_le16_clamp(unsigned char *dest, uint64_t n, int *clamped)\n \tcopy_le16(dest, clamp_max(n, 0xffff, clamped));\n }\n \n+static void copy_le32_clamp(unsigned char *dest, uint64_t n, int *clamped)\n+{\n+\tcopy_le32(dest, clamp_max(n, 0xffffffff, clamped));\n+}\n+\n static int strbuf_add_le(struct strbuf *sb, size_t size, uintmax_t n)\n {\n \twhile (size-- > 0) {\n@@ -154,6 +159,12 @@ static int strbuf_add_le(struct strbuf *sb, size_t size, uintmax_t n)\n \treturn -!!n;\n }\n \n+static uint32_t clamp32(uintmax_t n)\n+{\n+\tconst uintmax_t max = 0xffffffff;\n+\treturn (n < max) ? n : max;\n+}\n+\n static void *zlib_deflate_raw(void *data, unsigned long size,\n \t\t\t      int compression_level,\n \t\t\t      unsigned long *compressed_size)\n@@ -254,6 +265,8 @@ static int write_zip_entry(struct archiver_args *args,\n \tint is_binary = -1;\n \tconst char *path_without_prefix = path + args->baselen;\n \tunsigned int creator_version = 0;\n+\tsize_t zip_dir_extra_size = ZIP_EXTRA_MTIME_SIZE;\n+\tsize_t zip64_dir_extra_payload_size = 0;\n \n \tcrc = crc32(0, NULL, 0);\n \n@@ -433,6 +446,11 @@ static int write_zip_entry(struct archiver_args *args,\n \tfree(deflated);\n \tfree(buffer);\n \n+\tif (offset > 0xffffffff) {\n+\t\tzip64_dir_extra_payload_size += 8;\n+\t\tzip_dir_extra_size += 2 + 2 + zip64_dir_extra_payload_size;\n+\t}\n+\n \tstrbuf_add_le(&zip_dir, 4, 0x02014b50);\t/* magic */\n \tstrbuf_add_le(&zip_dir, 2, creator_version);\n \tstrbuf_add_le(&zip_dir, 2, 10);\t\t/* version */\n@@ -444,14 +462,20 @@ static int write_zip_entry(struct archiver_args *args,\n \tstrbuf_add_le(&zip_dir, 4, compressed_size);\n \tstrbuf_add_le(&zip_dir, 4, size);\n \tstrbuf_add_le(&zip_dir, 2, pathlen);\n-\tstrbuf_add_le(&zip_dir, 2, ZIP_EXTRA_MTIME_SIZE);\n+\tstrbuf_add_le(&zip_dir, 2, zip_dir_extra_size);\n \tstrbuf_add_le(&zip_dir, 2, 0);\t\t/* comment length */\n \tstrbuf_add_le(&zip_dir, 2, 0);\t\t/* disk */\n \tstrbuf_add_le(&zip_dir, 2, !is_binary);\n \tstrbuf_add_le(&zip_dir, 4, attr2);\n-\tstrbuf_add_le(&zip_dir, 4, offset);\n+\tstrbuf_add_le(&zip_dir, 4, clamp32(offset));\n \tstrbuf_add(&zip_dir, path, pathlen);\n \tstrbuf_add(&zip_dir, &extra, ZIP_EXTRA_MTIME_SIZE);\n+\tif (zip64_dir_extra_payload_size) {\n+\t\tstrbuf_add_le(&zip_dir, 2, 0x0001);\t/* magic */\n+\t\tstrbuf_add_le(&zip_dir, 2, zip64_dir_extra_payload_size);\n+\t\tif (offset >= 0xffffffff)\n+\t\t\tstrbuf_add_le(&zip_dir, 8, offset);\n+\t}\n \tzip_dir_entries++;\n \n \treturn 0;\n@@ -494,7 +518,7 @@ static void write_zip_trailer(const unsigned char *sha1)\n \t\t\t&clamped);\n \tcopy_le16_clamp(trailer.entries, zip_dir_entries, &clamped);\n \tcopy_le32(trailer.size, zip_dir.len);\n-\tcopy_le32(trailer.offset, zip_offset);\n+\tcopy_le32_clamp(trailer.offset, zip_offset, &clamped);\n \tcopy_le16(trailer.comment_length, sha1 ? GIT_SHA1_HEXSZ : 0);\n \n \twrite_or_die(1, zip_dir.buf, zip_dir.len);\ndiff --git a/t/t5004-archive-corner-cases.sh b/t/t5004-archive-corner-cases.sh\nindex bc052c803a..0ac94b5cc9 100755\n--- a/t/t5004-archive-corner-cases.sh\n+++ b/t/t5004-archive-corner-cases.sh\n@@ -155,7 +155,7 @@ test_expect_success ZIPINFO 'zip archive with many entries' '\n \ttest_cmp expect actual\n '\n \n-test_expect_failure EXPENSIVE,UNZIP 'zip archive bigger than 4GB' '\n+test_expect_success EXPENSIVE,UNZIP 'zip archive bigger than 4GB' '\n \t# build string containing 65536 characters\n \ts=0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef &&\n \ts=$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s &&\n-- \n2.12.2\n\n"},{"id":"317708","messageId":"e53f1d3f-be3a-ac28-89eb-63011da64586@web.de","threadId":"45770","inReplyTo":"85f2b6d1-107b-0624-af82-92446f28269e@web.de","subject":"[PATCH v3 5/5] archive-zip: support files bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-24T17:33:34Z","receivedAt":"2017-04-24T17:33:43Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Write a zip64 extended information extra field for big files as part of\ntheir local headers and as part of their central directory headers.\nAlso write a zip64 version of the data descriptor in that case.\n\nIf we're streaming then we don't know the compressed size at the time we\nwrite the header.  Deflate can end up making a file bigger instead of\nsmaller if we're unlucky.  Write a local zip64 header already for files\nwith a size of 2GB or more in this case to be on the safe side.\n\nBoth sizes need to be included in the local zip64 header, but the extra\nfield for the directory must only contain 64-bit equivalents for 32-bit\nvalues of 0xffffffff.\n\nSigned-off-by: Rene Scharfe <l.s.r@web.de>\n---\n archive-zip.c                   | 90 ++++++++++++++++++++++++++++++++++-------\n t/t5004-archive-corner-cases.sh |  2 +-\n 2 files changed, 76 insertions(+), 16 deletions(-)\n\ndiff --git a/archive-zip.c b/archive-zip.c\nindex 7d6f2a85d0..44ed78f163 100644\n--- a/archive-zip.c\n+++ b/archive-zip.c\n@@ -45,6 +45,14 @@ struct zip_data_desc {\n \tunsigned char _end[1];\n };\n \n+struct zip64_data_desc {\n+\tunsigned char magic[4];\n+\tunsigned char crc32[4];\n+\tunsigned char compressed_size[8];\n+\tunsigned char size[8];\n+\tunsigned char _end[1];\n+};\n+\n struct zip_dir_trailer {\n \tunsigned char magic[4];\n \tunsigned char disk[2];\n@@ -65,6 +73,14 @@ struct zip_extra_mtime {\n \tunsigned char _end[1];\n };\n \n+struct zip64_extra {\n+\tunsigned char magic[2];\n+\tunsigned char extra_size[2];\n+\tunsigned char size[8];\n+\tunsigned char compressed_size[8];\n+\tunsigned char _end[1];\n+};\n+\n struct zip64_dir_trailer {\n \tunsigned char magic[4];\n \tunsigned char record_size[8];\n@@ -94,11 +110,15 @@ struct zip64_dir_trailer_locator {\n  */\n #define ZIP_LOCAL_HEADER_SIZE\toffsetof(struct zip_local_header, _end)\n #define ZIP_DATA_DESC_SIZE\toffsetof(struct zip_data_desc, _end)\n+#define ZIP64_DATA_DESC_SIZE\toffsetof(struct zip64_data_desc, _end)\n #define ZIP_DIR_HEADER_SIZE\toffsetof(struct zip_dir_header, _end)\n #define ZIP_DIR_TRAILER_SIZE\toffsetof(struct zip_dir_trailer, _end)\n #define ZIP_EXTRA_MTIME_SIZE\toffsetof(struct zip_extra_mtime, _end)\n #define ZIP_EXTRA_MTIME_PAYLOAD_SIZE \\\n \t(ZIP_EXTRA_MTIME_SIZE - offsetof(struct zip_extra_mtime, flags))\n+#define ZIP64_EXTRA_SIZE\toffsetof(struct zip64_extra, _end)\n+#define ZIP64_EXTRA_PAYLOAD_SIZE \\\n+\t(ZIP64_EXTRA_SIZE - offsetof(struct zip64_extra, size))\n #define ZIP64_DIR_TRAILER_SIZE\toffsetof(struct zip64_dir_trailer, _end)\n #define ZIP64_DIR_TRAILER_RECORD_SIZE \\\n \t(ZIP64_DIR_TRAILER_SIZE - \\\n@@ -202,13 +222,23 @@ static void write_zip_data_desc(unsigned long size,\n \t\t\t\tunsigned long compressed_size,\n \t\t\t\tunsigned long crc)\n {\n-\tstruct zip_data_desc trailer;\n-\n-\tcopy_le32(trailer.magic, 0x08074b50);\n-\tcopy_le32(trailer.crc32, crc);\n-\tcopy_le32(trailer.compressed_size, compressed_size);\n-\tcopy_le32(trailer.size, size);\n-\twrite_or_die(1, &trailer, ZIP_DATA_DESC_SIZE);\n+\tif (size >= 0xffffffff || compressed_size >= 0xffffffff) {\n+\t\tstruct zip64_data_desc trailer;\n+\t\tcopy_le32(trailer.magic, 0x08074b50);\n+\t\tcopy_le32(trailer.crc32, crc);\n+\t\tcopy_le64(trailer.compressed_size, compressed_size);\n+\t\tcopy_le64(trailer.size, size);\n+\t\twrite_or_die(1, &trailer, ZIP64_DATA_DESC_SIZE);\n+\t\tzip_offset += ZIP64_DATA_DESC_SIZE;\n+\t} else {\n+\t\tstruct zip_data_desc trailer;\n+\t\tcopy_le32(trailer.magic, 0x08074b50);\n+\t\tcopy_le32(trailer.crc32, crc);\n+\t\tcopy_le32(trailer.compressed_size, compressed_size);\n+\t\tcopy_le32(trailer.size, size);\n+\t\twrite_or_die(1, &trailer, ZIP_DATA_DESC_SIZE);\n+\t\tzip_offset += ZIP_DATA_DESC_SIZE;\n+\t}\n }\n \n static void set_zip_header_data_desc(struct zip_local_header *header,\n@@ -252,6 +282,9 @@ static int write_zip_entry(struct archiver_args *args,\n \tstruct zip_local_header header;\n \tuintmax_t offset = zip_offset;\n \tstruct zip_extra_mtime extra;\n+\tstruct zip64_extra extra64;\n+\tsize_t header_extra_size = ZIP_EXTRA_MTIME_SIZE;\n+\tint need_zip64_extra = 0;\n \tunsigned long attr2;\n \tunsigned long compressed_size;\n \tunsigned long crc;\n@@ -344,21 +377,40 @@ static int write_zip_entry(struct archiver_args *args,\n \textra.flags[0] = 1;\t/* just mtime */\n \tcopy_le32(extra.mtime, args->time);\n \n+\tif (size > 0xffffffff || compressed_size > 0xffffffff)\n+\t\tneed_zip64_extra = 1;\n+\tif (stream && size > 0x7fffffff)\n+\t\tneed_zip64_extra = 1;\n+\n \tcopy_le32(header.magic, 0x04034b50);\n \tcopy_le16(header.version, 10);\n \tcopy_le16(header.flags, flags);\n \tcopy_le16(header.compression_method, method);\n \tcopy_le16(header.mtime, zip_time);\n \tcopy_le16(header.mdate, zip_date);\n-\tset_zip_header_data_desc(&header, size, compressed_size, crc);\n+\tif (need_zip64_extra) {\n+\t\tset_zip_header_data_desc(&header, 0xffffffff, 0xffffffff, crc);\n+\t\theader_extra_size += ZIP64_EXTRA_SIZE;\n+\t} else {\n+\t\tset_zip_header_data_desc(&header, size, compressed_size, crc);\n+\t}\n \tcopy_le16(header.filename_length, pathlen);\n-\tcopy_le16(header.extra_length, ZIP_EXTRA_MTIME_SIZE);\n+\tcopy_le16(header.extra_length, header_extra_size);\n \twrite_or_die(1, &header, ZIP_LOCAL_HEADER_SIZE);\n \tzip_offset += ZIP_LOCAL_HEADER_SIZE;\n \twrite_or_die(1, path, pathlen);\n \tzip_offset += pathlen;\n \twrite_or_die(1, &extra, ZIP_EXTRA_MTIME_SIZE);\n \tzip_offset += ZIP_EXTRA_MTIME_SIZE;\n+\tif (need_zip64_extra) {\n+\t\tcopy_le16(extra64.magic, 0x0001);\n+\t\tcopy_le16(extra64.extra_size, ZIP64_EXTRA_PAYLOAD_SIZE);\n+\t\tcopy_le64(extra64.size, size);\n+\t\tcopy_le64(extra64.compressed_size, compressed_size);\n+\t\twrite_or_die(1, &extra64, ZIP64_EXTRA_SIZE);\n+\t\tzip_offset += ZIP64_EXTRA_SIZE;\n+\t}\n+\n \tif (stream && method == 0) {\n \t\tunsigned char buf[STREAM_BUFFER_SIZE];\n \t\tssize_t readlen;\n@@ -381,7 +433,6 @@ static int write_zip_entry(struct archiver_args *args,\n \t\tzip_offset += compressed_size;\n \n \t\twrite_zip_data_desc(size, compressed_size, crc);\n-\t\tzip_offset += ZIP_DATA_DESC_SIZE;\n \t} else if (stream && method == 8) {\n \t\tunsigned char buf[STREAM_BUFFER_SIZE];\n \t\tssize_t readlen;\n@@ -437,7 +488,6 @@ static int write_zip_entry(struct archiver_args *args,\n \t\tzip_offset += compressed_size;\n \n \t\twrite_zip_data_desc(size, compressed_size, crc);\n-\t\tzip_offset += ZIP_DATA_DESC_SIZE;\n \t} else if (compressed_size > 0) {\n \t\twrite_or_die(1, out, compressed_size);\n \t\tzip_offset += compressed_size;\n@@ -446,8 +496,14 @@ static int write_zip_entry(struct archiver_args *args,\n \tfree(deflated);\n \tfree(buffer);\n \n-\tif (offset > 0xffffffff) {\n-\t\tzip64_dir_extra_payload_size += 8;\n+\tif (compressed_size > 0xffffffff || size > 0xffffffff ||\n+\t    offset > 0xffffffff) {\n+\t\tif (compressed_size >= 0xffffffff)\n+\t\t\tzip64_dir_extra_payload_size += 8;\n+\t\tif (size >= 0xffffffff)\n+\t\t\tzip64_dir_extra_payload_size += 8;\n+\t\tif (offset >= 0xffffffff)\n+\t\t\tzip64_dir_extra_payload_size += 8;\n \t\tzip_dir_extra_size += 2 + 2 + zip64_dir_extra_payload_size;\n \t}\n \n@@ -459,8 +515,8 @@ static int write_zip_entry(struct archiver_args *args,\n \tstrbuf_add_le(&zip_dir, 2, zip_time);\n \tstrbuf_add_le(&zip_dir, 2, zip_date);\n \tstrbuf_add_le(&zip_dir, 4, crc);\n-\tstrbuf_add_le(&zip_dir, 4, compressed_size);\n-\tstrbuf_add_le(&zip_dir, 4, size);\n+\tstrbuf_add_le(&zip_dir, 4, clamp32(compressed_size));\n+\tstrbuf_add_le(&zip_dir, 4, clamp32(size));\n \tstrbuf_add_le(&zip_dir, 2, pathlen);\n \tstrbuf_add_le(&zip_dir, 2, zip_dir_extra_size);\n \tstrbuf_add_le(&zip_dir, 2, 0);\t\t/* comment length */\n@@ -473,6 +529,10 @@ static int write_zip_entry(struct archiver_args *args,\n \tif (zip64_dir_extra_payload_size) {\n \t\tstrbuf_add_le(&zip_dir, 2, 0x0001);\t/* magic */\n \t\tstrbuf_add_le(&zip_dir, 2, zip64_dir_extra_payload_size);\n+\t\tif (size >= 0xffffffff)\n+\t\t\tstrbuf_add_le(&zip_dir, 8, size);\n+\t\tif (compressed_size >= 0xffffffff)\n+\t\t\tstrbuf_add_le(&zip_dir, 8, compressed_size);\n \t\tif (offset >= 0xffffffff)\n \t\t\tstrbuf_add_le(&zip_dir, 8, offset);\n \t}\ndiff --git a/t/t5004-archive-corner-cases.sh b/t/t5004-archive-corner-cases.sh\nindex 0ac94b5cc9..a6875dfdb1 100755\n--- a/t/t5004-archive-corner-cases.sh\n+++ b/t/t5004-archive-corner-cases.sh\n@@ -178,7 +178,7 @@ test_expect_success EXPENSIVE,UNZIP 'zip archive bigger than 4GB' '\n \t\"$GIT_UNZIP\" -t many-big.zip\n '\n \n-test_expect_failure EXPENSIVE,UNZIP 'zip archive with files bigger than 4GB' '\n+test_expect_success EXPENSIVE,UNZIP,ZIPINFO 'zip archive with files bigger than 4GB' '\n \t# Pack created with:\n \t#   dd if=/dev/zero of=file bs=1M count=4100 && git hash-object -w file\n \tmkdir -p .git/objects/pack &&\n-- \n2.12.2\n\n"},{"id":"317711","messageId":"alpine.DEB.2.11.1704241912510.30460@perkele.intern.softwolves.pp.se","threadId":"45770","inReplyTo":"02ddca3c-a11f-7c0c-947e-5ca87a62cdee@web.de","subject":"Re: [PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2017-04-24T18:24:53Z","receivedAt":"2017-04-24T18:25:13Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"René Scharfe:\n\n> @@ -433,6 +446,11 @@ static int write_zip_entry(struct archiver_args *args,\n> \tfree(deflated);\n> \tfree(buffer);\n>\n> +\tif (offset > 0xffffffff) {\n> +\t\tzip64_dir_extra_payload_size += 8;\n> +\t\tzip_dir_extra_size += 2 + 2 + zip64_dir_extra_payload_size;\n> +\t}\n> +\n> \tstrbuf_add_le(&zip_dir, 4, 0x02014b50);\t/* magic */\n> \tstrbuf_add_le(&zip_dir, 2, creator_version);\n> \tstrbuf_add_le(&zip_dir, 2, 10);\t\t/* version */\n\nThis needs to be >=. The spec says that if the value is 0xffffffff, \nthere should be a zip64 record with the actual size (even if it is \n0xffffffff).\n\nAlso set the version required to 45 (4.5) for any record that has zip64 \nfields.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"317715","messageId":"d453610f-dbd5-3f6c-d386-69a74c238b11@web.de","threadId":"45770","inReplyTo":"alpine.DEB.2.11.1704241912510.30460@perkele.intern.softwolves.pp.se","subject":"Re: [PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-24T20:06:11Z","receivedAt":"2017-04-24T20:06:23Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 24.04.2017 um 20:24 schrieb Peter Krefting:\n> René Scharfe:\n> \n>> @@ -433,6 +446,11 @@ static int write_zip_entry(struct archiver_args \n>> *args,\n>>     free(deflated);\n>>     free(buffer);\n>>\n>> +    if (offset > 0xffffffff) {\n>> +        zip64_dir_extra_payload_size += 8;\n>> +        zip_dir_extra_size += 2 + 2 + zip64_dir_extra_payload_size;\n>> +    }\n>> +\n>>     strbuf_add_le(&zip_dir, 4, 0x02014b50);    /* magic */\n>>     strbuf_add_le(&zip_dir, 2, creator_version);\n>>     strbuf_add_le(&zip_dir, 2, 10);        /* version */\n> \n> This needs to be >=. The spec says that if the value is 0xffffffff, \n> there should be a zip64 record with the actual size (even if it is \n> 0xffffffff).\n\nCould you please cite the relevant part?\n\nHere's how I read it: If a value doesn't fit into a 32-bit field it is \nset to 0xffffffff, a zip64 extra is added and a 64-bit field stores the \nactual value.  The magic value 0xffffffff indicates that a corresponding \n64-bit field is present in the zip64 extra.  That means even if a value \nis 0xffffffff (and thus fits) we need to add it to the zip64 extra.  If \nthere is no zip64 extra then we can store 0xffffffff in the 32-bit \nfield, though.\n\n> Also set the version required to 45 (4.5) for any record that has zip64 \n> fields.\n\nAh, yes indeed.\n\nRené\n"},{"id":"317718","messageId":"b9945842-7c34-fd02-0178-970613b242f6@web.de","threadId":"45770","inReplyTo":"d453610f-dbd5-3f6c-d386-69a74c238b11@web.de","subject":"Re: [PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-24T20:39:25Z","receivedAt":"2017-04-24T20:39:35Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 24.04.2017 um 22:06 schrieb René Scharfe:\n> Am 24.04.2017 um 20:24 schrieb Peter Krefting:\n>> René Scharfe:\n>> Also set the version required to 45 (4.5) for any record that has \n>> zip64 fields.\n> \n> Ah, yes indeed.\n\nWhen I tried to implement this I realized that should set 20 for \ndirectories, but we use 10 for everything.  Then I compared with InfoZIP \nzip and noticed that they only ever sets 45 for the zip64 end of central \ndirectory record and 10 for everything else.  So I think we can/should \nkeep it as it is, for compatibility's sake -- unless there is a problem \nwith an old (but still relevant) extractor.\n\nRené\n"},{"id":"317720","messageId":"13bc4fd9-984d-8b37-a18a-c7d273fbba36@kdbg.org","threadId":"45770","inReplyTo":"d453610f-dbd5-3f6c-d386-69a74c238b11@web.de","subject":"Re: [PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2017-04-24T21:02:14Z","receivedAt":"2017-04-24T21:02:26Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 24.04.2017 um 22:06 schrieb René Scharfe:\n> Am 24.04.2017 um 20:24 schrieb Peter Krefting:\n>> René Scharfe:\n>>\n>>> @@ -433,6 +446,11 @@ static int write_zip_entry(struct archiver_args\n>>> *args,\n>>>     free(deflated);\n>>>     free(buffer);\n>>>\n>>> +    if (offset > 0xffffffff) {\n>>> +        zip64_dir_extra_payload_size += 8;\n>>> +        zip_dir_extra_size += 2 + 2 + zip64_dir_extra_payload_size;\n>>> +    }\n>>> +\n>>>     strbuf_add_le(&zip_dir, 4, 0x02014b50);    /* magic */\n>>>     strbuf_add_le(&zip_dir, 2, creator_version);\n>>>     strbuf_add_le(&zip_dir, 2, 10);        /* version */\n>>\n>> This needs to be >=. The spec says that if the value is 0xffffffff,\n>> there should be a zip64 record with the actual size (even if it is\n>> 0xffffffff).\n>\n> Could you please cite the relevant part?\n>\n> Here's how I read it: If a value doesn't fit into a 32-bit field it is\n> set to 0xffffffff, a zip64 extra is added and a 64-bit field stores the\n> actual value.  The magic value 0xffffffff indicates that a corresponding\n> 64-bit field is present in the zip64 extra.  That means even if a value\n> is 0xffffffff (and thus fits) we need to add it to the zip64 extra.  If\n> there is no zip64 extra then we can store 0xffffffff in the 32-bit\n> field, though.\n\nThe reader I wrote recently interprets 0xffffffff as special if the \nversion is 45. Then, if there is no zip64 extra record, it is a broken \nZIP archive. You are saying that my reader is wrong in this special case...\n\n-- Hannes\n\n"},{"id":"317721","messageId":"0E87DF4B-1924-41DE-97E0-8294146E7D4E@blackthorn-media.com","threadId":"45770","inReplyTo":"e53f1d3f-be3a-ac28-89eb-63011da64586@web.de","subject":"Re: [PATCH v3 5/5] archive-zip: support files bigger than 4GB","fromName":"Keith Goldfarb","fromEmail":"keith@blackthorn-media.com","sentAt":"2017-04-24T21:11:31Z","receivedAt":"2017-04-24T21:11:39Z","isPatch":true,"sender":{"key":"keith@blackthorn-media.com","avatar":null},"body":"This set of patches works for my test case.\n\nThanks,\n\nK.\n\n"},{"id":"317726","messageId":"5dcbd1a6-056b-9a81-746d-329d5dafad3b@web.de","threadId":"45770","inReplyTo":"13bc4fd9-984d-8b37-a18a-c7d273fbba36@kdbg.org","subject":"Re: [PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-24T21:41:31Z","receivedAt":"2017-04-24T21:41:41Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 24.04.2017 um 23:02 schrieb Johannes Sixt:\n> Am 24.04.2017 um 22:06 schrieb René Scharfe:\n>> Am 24.04.2017 um 20:24 schrieb Peter Krefting:\n>>> René Scharfe:\n>>>\n>>>> @@ -433,6 +446,11 @@ static int write_zip_entry(struct archiver_args\n>>>> *args,\n>>>>     free(deflated);\n>>>>     free(buffer);\n>>>>\n>>>> +    if (offset > 0xffffffff) {\n>>>> +        zip64_dir_extra_payload_size += 8;\n>>>> +        zip_dir_extra_size += 2 + 2 + zip64_dir_extra_payload_size;\n>>>> +    }\n>>>> +\n>>>>     strbuf_add_le(&zip_dir, 4, 0x02014b50);    /* magic */\n>>>>     strbuf_add_le(&zip_dir, 2, creator_version);\n>>>>     strbuf_add_le(&zip_dir, 2, 10);        /* version */\n>>>\n>>> This needs to be >=. The spec says that if the value is 0xffffffff,\n>>> there should be a zip64 record with the actual size (even if it is\n>>> 0xffffffff).\n>>\n>> Could you please cite the relevant part?\n>>\n>> Here's how I read it: If a value doesn't fit into a 32-bit field it is\n>> set to 0xffffffff, a zip64 extra is added and a 64-bit field stores the\n>> actual value.  The magic value 0xffffffff indicates that a corresponding\n>> 64-bit field is present in the zip64 extra.  That means even if a value\n>> is 0xffffffff (and thus fits) we need to add it to the zip64 extra.  If\n>> there is no zip64 extra then we can store 0xffffffff in the 32-bit\n>> field, though.\n> \n> The reader I wrote recently interprets 0xffffffff as special if the \n> version is 45. Then, if there is no zip64 extra record, it is a broken \n> ZIP archive. You are saying that my reader is wrong in this special case...\n\nThe \"version needed to extract\" field is not particularly well-suited\nfor feature detection.  If you e.g. compress with LZMA then it should be\nset to 63 even for small files that need no zip64 extra field.  So it\nseems to rather be meant to allow readers to spot headers they can't\ninterpret because they were written against an earlier version of the\nspec.\n\nInfoZIP's zip writes 10 into the version field, no matter if the entry\nhas a zip64 extra field or not, as I wrote in my other reply.  That\nseems to collide with the spec, but I'd expect readers to be able to\nhandle its output.\n\nRené\n"},{"id":"317774","messageId":"xmqq7f29w0uj.fsf@gitster.mtv.corp.google.com","threadId":"45770","inReplyTo":"e53f1d3f-be3a-ac28-89eb-63011da64586@web.de","subject":"Re: [PATCH v3 5/5] archive-zip: support files bigger than 4GB","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-04-25T04:46:44Z","receivedAt":"2017-04-25T04:46:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> diff --git a/t/t5004-archive-corner-cases.sh b/t/t5004-archive-corner-cases.sh\n> index 0ac94b5cc9..a6875dfdb1 100755\n> --- a/t/t5004-archive-corner-cases.sh\n> +++ b/t/t5004-archive-corner-cases.sh\n> @@ -178,7 +178,7 @@ test_expect_success EXPENSIVE,UNZIP 'zip archive bigger than 4GB' '\n>  \t\"$GIT_UNZIP\" -t many-big.zip\n>  '\n>  \n> -test_expect_failure EXPENSIVE,UNZIP 'zip archive with files bigger than 4GB' '\n> +test_expect_success EXPENSIVE,UNZIP,ZIPINFO 'zip archive with files bigger than 4GB' '\n\nThis is a bit curious, as 1/5 adds this test that expects a failure\nwith three prerequisites already.  I'll assume that this is a rebase\nglitch and the preimage actually must have ,ZIPINFO there already.\n\n>  \t# Pack created with:\n>  \t#   dd if=/dev/zero of=file bs=1M count=4100 && git hash-object -w file\n>  \tmkdir -p .git/objects/pack &&\n"},{"id":"317775","messageId":"xmqq37cxw0mf.fsf@gitster.mtv.corp.google.com","threadId":"45770","inReplyTo":"949f19e6-0414-9abc-9754-064d7e58c169@web.de","subject":"Re: [PATCH v3 2/5] archive-zip: use strbuf for ZIP directory","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-04-25T04:51:36Z","receivedAt":"2017-04-25T04:51:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> Keep the ZIP central director, which is written after all archive\n\nIs this typoed \"directorY\"?  I know there was discussion on the\ncorrect terminology I only skimmed and I am too lazy to go back\nthere to understand it to answer this question myself, so...\n\n> -#define ZIP_DIRECTORY_MIN_SIZE\t(1024 * 1024)\n\nThis tells me that there is this thing called ZIP directory, so\nprobably my guess above is correct ;-)\n\n"},{"id":"317776","messageId":"80415bd2-1669-4dcc-4c08-87c096f9b063@web.de","threadId":"45770","inReplyTo":"xmqq37cxw0mf.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v3 2/5] archive-zip: use strbuf for ZIP directory","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-25T05:28:00Z","receivedAt":"2017-04-25T05:28:14Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 25.04.2017 um 06:51 schrieb Junio C Hamano:\n> René Scharfe <l.s.r@web.de> writes:\n> \n>> Keep the ZIP central director, which is written after all archive\n> \n> Is this typoed \"directorY\"?  I know there was discussion on the\n> correct terminology I only skimmed and I am too lazy to go back\n> there to understand it to answer this question myself, so...\n\nYep, a typo, this patch touches a \"directory\", as stated in the subject.\n\nThe \"ZIP central director\" would be PKWARE Inc., I guess? ;-)  The patch \ndoesn't affect them..\n\nRené\n"},{"id":"317777","messageId":"10e59c7b-d8e4-e941-893a-bf4a68a54da0@web.de","threadId":"45770","inReplyTo":"xmqq7f29w0uj.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v3 5/5] archive-zip: support files bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-25T05:27:56Z","receivedAt":"2017-04-25T05:28:16Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 25.04.2017 um 06:46 schrieb Junio C Hamano:\n> René Scharfe <l.s.r@web.de> writes:\n> \n>> diff --git a/t/t5004-archive-corner-cases.sh b/t/t5004-archive-corner-cases.sh\n>> index 0ac94b5cc9..a6875dfdb1 100755\n>> --- a/t/t5004-archive-corner-cases.sh\n>> +++ b/t/t5004-archive-corner-cases.sh\n>> @@ -178,7 +178,7 @@ test_expect_success EXPENSIVE,UNZIP 'zip archive bigger than 4GB' '\n>>   \t\"$GIT_UNZIP\" -t many-big.zip\n>>   '\n>>   \n>> -test_expect_failure EXPENSIVE,UNZIP 'zip archive with files bigger than 4GB' '\n>> +test_expect_success EXPENSIVE,UNZIP,ZIPINFO 'zip archive with files bigger than 4GB' '\n> \n> This is a bit curious, as 1/5 adds this test that expects a failure\n> with three prerequisites already.  I'll assume that this is a rebase\n> glitch and the preimage actually must have ,ZIPINFO there already.\n\nYes, indeed -- I shouldn't try to do last minute edits.\n\nRené\n"},{"id":"317787","messageId":"alpine.DEB.2.11.1704250851420.23677@perkele.intern.softwolves.pp.se","threadId":"45770","inReplyTo":"d453610f-dbd5-3f6c-d386-69a74c238b11@web.de","subject":"Re: [PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2017-04-25T07:55:50Z","receivedAt":"2017-04-25T07:55:57Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"René Scharfe:\n\n>> This needs to be >=. The spec says that if the value is 0xffffffff, there \n>> should be a zip64 record with the actual size (even if it is 0xffffffff).\n> Could you please cite the relevant part?\n\n4.4.8 compressed size: (4 bytes)\n4.4.9 uncompressed size: (4 bytes)\n\n\"If an archive is in ZIP64 format and the value in this field is \n0xFFFFFFFF, the size will be in the corresponding 8 byte ZIP64 \nextended information extra field.\"\n\n\nOf course, there is no definition of how they define that \"an archive \nis in ZIP64 format\", but I would say that is whenever it has any ZIP64 \nstructures.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"317818","messageId":"fdc17512-94dc-4f7f-4fd3-f933e1b18e8f@web.de","threadId":"45770","inReplyTo":"alpine.DEB.2.11.1704250851420.23677@perkele.intern.softwolves.pp.se","subject":"Re: [PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-25T16:24:47Z","receivedAt":"2017-04-25T16:25:11Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 25.04.2017 um 09:55 schrieb Peter Krefting:\n> René Scharfe:\n> \n>>> This needs to be >=. The spec says that if the value is 0xffffffff, \n>>> there should be a zip64 record with the actual size (even if it is \n>>> 0xffffffff).\n>> Could you please cite the relevant part?\n> \n> 4.4.8 compressed size: (4 bytes)\n> 4.4.9 uncompressed size: (4 bytes)\n> \n> \"If an archive is in ZIP64 format and the value in this field is \n> 0xFFFFFFFF, the size will be in the corresponding 8 byte ZIP64 extended \n> information extra field.\"\n> \n> \n> Of course, there is no definition of how they define that \"an archive is \n> in ZIP64 format\", but I would say that is whenever it has any ZIP64 \n> structures.\n\nI struggled with that sentence as well.  There is no explicit \"format\"\nfield AFAICS.  The closest at the archive level are zip64 end of central\ndirectory record and locator.  But what really matters is the presence\nof a zip64 extended information extra field to hold the 64-bit size\nvalue.\n\nThere's also this general note a bit higher up:\n\n       \"4.4.1.4  If one of the fields in the end of central directory\n       record is too small to hold required data, the field should be\n       set to -1 (0xFFFF or 0xFFFFFFFF) and the ZIP64 format record\n       should be created.\"\n\nMy interpretation: An archiver that can only emit 32-bit ZIP files\n(either because it doesn't support ZIP64 or due to a compatibility\noption set by the user) writes 32-bit size fields and has no defined way\nto deal with overflows.  An archiver that is allowed to use ZIP64 can\nemit zip64 extras as needed.\n\nOr in other words: A legacy ZIP archive and a ZIP64 archive can be\nbit-wise the same if all values for all entries fit into the legacy\nfields, but the difference in terms of the spec is what the archiver was\nallowed to do when it created them.\n\n\t# 4-byte sizes, not ZIP64\n\tarch --format=zip ...\n\n\t# ZIP64, can use 8-byte sizes as needed\n\tarch --format=zip64 ...\n\nMakes sense?\n\nRené\n"},{"id":"317998","messageId":"alpine.DEB.2.11.1704262154420.29054@perkele.intern.softwolves.pp.se","threadId":"45770","inReplyTo":"fdc17512-94dc-4f7f-4fd3-f933e1b18e8f@web.de","subject":"Re: [PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2017-04-26T21:02:33Z","receivedAt":"2017-04-26T21:02:41Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"René Scharfe:\n\n> I struggled with that sentence as well.  There is no explicit \n> \"format\" field AFAICS.\n\nExactly. I interpret that as it is in zip64 format if there are any \nzip64 structures in the archive (especially if there is a zip64 \nend of central directory locator).\n\n> Or in other words: A legacy ZIP archive and a ZIP64 archive can be \n> bit-wise the same if all values for all entries fit into the legacy \n> fields, but the difference in terms of the spec is what the archiver \n> was allowed to do when it created them.\n\nAs long as all sizes are below (unsigned) -1, then they would be \nidentical. If one, and only one, of the sizes are equal to (unsigned) \n-1 (and none overflow), then it is up to intepretation whether or not \na ZIP64-aware archiver is allowed to output an archive that is not in \nZIP64 format. If any single size or value overflows the 32 (16) bit \nvalues, then ZIP64 format is needed.\n\n> \t# 4-byte sizes, not ZIP64\n> \tarch --format=zip ...\n>\n> \t# ZIP64, can use 8-byte sizes as needed\n> \tarch --format=zip64 ...\n>\n> Makes sense?\n\nWell, I would say that it would be a lot easier to always emit zip64 \narchives. An old-style unzipper should be able to read them anyway if \nthere are no overflowing fields, right? And, besides, who in 2017 has \nan unzip tool that is unable to read zip64? Info-Zip UnZip has \nsupported Zip64 since 2009.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"318014","messageId":"87470c8c-e061-e4b3-42fe-84a30858fc0d@web.de","threadId":"45770","inReplyTo":"alpine.DEB.2.11.1704262154420.29054@perkele.intern.softwolves.pp.se","subject":"Re: [PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-26T23:38:06Z","receivedAt":"2017-04-26T23:38:19Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 26.04.2017 um 23:02 schrieb Peter Krefting:\n> René Scharfe:\n> \n>> I struggled with that sentence as well.  There is no explicit \"format\" \n>> field AFAICS.\n> \n> Exactly. I interpret that as it is in zip64 format if there are any \n> zip64 structures in the archive (especially if there is a zip64 end of \n> central directory locator).\n\nThe crucial point is that I think the choice is per entry, i.e. if we\nhad to write a zip64 record for one file we can still emit a legacy\nrecord for the next file that has a size of 0xffffffff.\n\n>> Or in other words: A legacy ZIP archive and a ZIP64 archive can be \n>> bit-wise the same if all values for all entries fit into the legacy \n>> fields, but the difference in terms of the spec is what the archiver \n>> was allowed to do when it created them.\n> \n> As long as all sizes are below (unsigned) -1, then they would be \n> identical. If one, and only one, of the sizes are equal to (unsigned) -1 \n> (and none overflow), then it is up to intepretation whether or not a \n> ZIP64-aware archiver is allowed to output an archive that is not in \n> ZIP64 format. If any single size or value overflows the 32 (16) bit \n> values, then ZIP64 format is needed.\n\nSizes can be stored in zip64 entries even if they are lower (from a\nparagraph about the data descriptor):\n\n\"4.3.9.2 When compressing files, compressed and uncompressed sizes\n       should be stored in ZIP64 format (as 8 byte values) when a\n       file's size exceeds 0xFFFFFFFF.   However ZIP64 format may be\n       used regardless of the size of a file.\"\n\n(But I don't see a benefit.)\n\n>>     # 4-byte sizes, not ZIP64\n>>     arch --format=zip ...\n>>\n>>     # ZIP64, can use 8-byte sizes as needed\n>>     arch --format=zip64 ...\n>>\n>> Makes sense?\n> \n> Well, I would say that it would be a lot easier to always emit zip64 \n> archives. An old-style unzipper should be able to read them anyway if \n> there are no overflowing fields, right? And, besides, who in 2017 has an \n> unzip tool that is unable to read zip64? Info-Zip UnZip has supported \n> Zip64 since 2009.\n\nWindows XP.  Don't laugh. ;)\n\nIf you write zip64 extras for all size records then an old extractor\nwill only see the value 0xffffffff in them and ignore the zip64 part --\nor ignore the entries outright.\n\nWriting zip64 records only as needed saves space -- and that's what\nzipping is all about, isn't it?\n\nAdding unnecessary zip64 records would produce different ZIP files than\nearlier version of git archive.  That's not a strong argument as changes\nto libz can potentially do the same, but it still might affect someone\nwho caches generated ZIP files.\n\nWhat I sent matches the behavior of InfoZIP zip (modulo bugs).  Why not\nfollow their lead?\n\n(And one of the bugs in my patches not setting the version field to 45\nas you pointed out earlier already.  InfoZIP may forget to do that if it\nuses a zip64 extra for recording the offset, but it does set the version\ncorrectly for files bigger than 4GB.)\n\nWhat do other archivers do?\n\nBut I think a more important question is: Can the generated files\nbe extracted by popular tools (most importantly Windows' built-in\nfunctionality, I guess)?\n\nRené\n"},{"id":"318038","messageId":"alpine.DEB.2.11.1704270552590.4681@perkele.intern.softwolves.pp.se","threadId":"45770","inReplyTo":"87470c8c-e061-e4b3-42fe-84a30858fc0d@web.de","subject":"Re: [PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2017-04-27T04:57:32Z","receivedAt":"2017-04-27T04:57:43Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"René Scharfe:\n\n> Sizes can be stored in zip64 entries even if they are lower (from a \n> paragraph about the data descriptor):\n>\n> \"4.3.9.2 When compressing files, compressed and uncompressed sizes\n>      should be stored in ZIP64 format (as 8 byte values) when a\n>      file's size exceeds 0xFFFFFFFF.   However ZIP64 format may be\n>      used regardless of the size of a file.\"\n\nThat is only for the data descriptor. And this is the really confusing \none as it only has the sizes, and do not have an indication anywhere \nof which format it is in. So they are either 32-bit or 64-bit, \ndepending on whether it is a ZIP64 format archive, but it doesn't \ndefine how to tell.\n\nFor regular entries, the text is a bit clearer about the -1, though.\n\n> Windows XP.  Don't laugh. ;)\n\nYou can always install 7-zip or something to extract on XP.\n\n> What do other archivers do?\n\nYou should compare with what bsdtar (libarchive) does in zip64 mode. \nIt also only ever does streaming mode (with data descriptors and \nsuch), and it does zip64.\n\n> But I think a more important question is: Can the generated files be \n> extracted by popular tools (most importantly Windows' built-in \n> functionality, I guess)?\n\nOK, so only enable zip64 mode if there are files >4G or the archive \nends up being >4G. But the question is how we can tell, especially in \nstreaming mode, and especially if data descriptors are magical...\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"318085","messageId":"a9cb6572-500e-bbc6-2aac-7cb940d4b171@web.de","threadId":"45770","inReplyTo":"alpine.DEB.2.11.1704270552590.4681@perkele.intern.softwolves.pp.se","subject":"Re: [PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-27T19:54:52Z","receivedAt":"2017-04-27T19:55:19Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 27.04.2017 um 06:57 schrieb Peter Krefting:\n> René Scharfe:\n>> Windows XP.  Don't laugh. ;)\n>\n> You can always install 7-zip or something to extract on XP.\n\nSure, but if we were to start emitting zip64 records regardless of the\nsize of entries then we'd break compatibility.  We should have a very\ngood reason for doing that.  (I don't see the need so far.)\n\n>> What do other archivers do?\n> \n> You should compare with what bsdtar (libarchive) does in zip64 mode. It \n> also only ever does streaming mode (with data descriptors and such), and \n> it does zip64.\n\nGood idea.  They do the same as InfoZIP (whose behavior I copied in\nthe series), i.e. emit zip64 records only for big files by default:\n\nhttps://github.com/libarchive/libarchive/blob/master/libarchive/archive_write_set_format_zip.c#L722\n\nThey set the bar much higher for the uncompressed size of streamed\nfiles; they emit zip64 extras only for sizes bigger than 0xff000000.\n\n>> But I think a more important question is: Can the generated files be \n>> extracted by popular tools (most importantly Windows' built-in \n>> functionality, I guess)?\n> \n> OK, so only enable zip64 mode if there are files >4G or the archive ends \n> up being >4G. But the question is how we can tell, especially in \n> streaming mode, and especially if data descriptors are magical...\n\nThe type of descriptor to use depends on the presence of 64-bit\nsizes in a zip64 extra for that record.  For streaming compression\nsome kind of threshold lower than 0xffffffff needs to be set,\nbecause deflate can increase the size of the result.\n\nRené\n"},{"id":"318127","messageId":"alpine.DEB.2.00.1704280936290.1440@ds9.cixit.se","threadId":"45770","inReplyTo":"a9cb6572-500e-bbc6-2aac-7cb940d4b171@web.de","subject":"Re: [PATCH v3 4/5] archive-zip: support archives bigger than 4GB","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2017-04-28T08:40:09Z","receivedAt":"2017-04-28T08:41:28Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"René Scharfe:\n\n> Sure, but if we were to start emitting zip64 records regardless of \n> the size of entries then we'd break compatibility.  We should have a \n> very good reason for doing that.  (I don't see the need so far.)\n\nSure, sounds good.\n\n> The type of descriptor to use depends on the presence of 64-bit \n> sizes in a zip64 extra for that record.  For streaming compression \n> some kind of threshold lower than 0xffffffff needs to be set, \n> because deflate can increase the size of the result.\n\nIndeed. And it seems that they use the version identifier (>= 45) to \ncheck whether it is in \"zip64 format\" or not. It seems a bit hit or \nmiss to me, the best would be to always use the pre-amble descriptor, \nbut that requires holding the entire compressed data in memory (or \nusing temporary files or running two passes), neither which are very \ngood ideas.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"318246","messageId":"3df2b03f-ab86-09ac-0fc8-3c6eb10c6704@web.de","threadId":"45770","inReplyTo":"85f2b6d1-107b-0624-af82-92446f28269e@web.de","subject":"Re: [PATCH v3 0/5] archive-zip: support files and archives bigger than 4GB","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2017-04-29T21:00:52Z","receivedAt":"2017-04-29T21:01:14Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 2017-04-24 19:22, René Scharfe wrote:\n> The first patch adds (expensive) tests, the next two are cleanups which\n> set the stage for the remaining two to actually implement zip64 support\n> for offsets and file sizes.\n> \n> Half of the series had been laying around for months, half-finished and\n> forgotten because I got distracted by the holiday season. :-/\n> \n>   archive-zip: add tests for big ZIP archives\n>   archive-zip: use strbuf for ZIP directory\n>   archive-zip: write ZIP dir entry directly to strbuf\n>   archive-zip: support archives bigger than 4GB\n>   archive-zip: support files bigger than 4GB\n> \n>  archive-zip.c                   | 211 ++++++++++++++++++++++++----------------\n>  t/t5004-archive-corner-cases.sh |  45 +++++++++\n>  t/t5004/big-pack.zip            | Bin 0 -> 7373 bytes\n>  3 files changed, 172 insertions(+), 84 deletions(-)\n>  create mode 100644 t/t5004/big-pack.zip\n> \nThis fails here under Mac OS:\ncommit 4cdf3f9d84568da72f1dcade812de7a42ecb6d15\nAuthor: René Scharfe <l.s.r@web.de>\nDate:   Mon Apr 24 19:33:34 2017 +0200\n\n    archive-zip: support files bigger than 4GB\n\n---------------------------\nParts of t5004.log, hope this is helpful:\n\n\"$GIT_UNZIP\" -t many-big.zip\n\nArchive:  many-big.zip\nwarning [many-big.zip]:  577175 extra bytes at beginning or within zipfile\n  (attempting to process anyway)\nerror [many-big.zip]:  start of central directory not found;\n  zipfile corrupt.\n  (please check that you have transferred or created the zipfile in the\n  appropriate BINARY mode and that you have compiled UnZip properly)\nnot ok 12 - zip archive bigger than 4GB\n#\t\n#\t\t# build string containing 65536 characters\n\n"},{"id":"318248","messageId":"edf33657-f74b-3cd5-44a7-8e16231bd978@web.de","threadId":"45770","inReplyTo":"3df2b03f-ab86-09ac-0fc8-3c6eb10c6704@web.de","subject":"Re: [PATCH v3 0/5] archive-zip: support files and archives bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-29T22:28:17Z","receivedAt":"2017-04-29T22:28:29Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 29.04.2017 um 23:00 schrieb Torsten Bögershausen:\n> This fails here under Mac OS:\n> commit 4cdf3f9d84568da72f1dcade812de7a42ecb6d15\n> Author: René Scharfe <l.s.r@web.de>\n> Date:   Mon Apr 24 19:33:34 2017 +0200\n> \n>      archive-zip: support files bigger than 4GB\n> \n> ---------------------------\n> Parts of t5004.log, hope this is helpful:\n> \n> \"$GIT_UNZIP\" -t many-big.zip\n> \n> Archive:  many-big.zip\n> warning [many-big.zip]:  577175 extra bytes at beginning or within zipfile\n>    (attempting to process anyway)\n> error [many-big.zip]:  start of central directory not found;\n>    zipfile corrupt.\n>    (please check that you have transferred or created the zipfile in the\n>    appropriate BINARY mode and that you have compiled UnZip properly)\n> not ok 12 - zip archive bigger than 4GB\n> #\t\n> #\t\t# build string containing 65536 characters\n\nWhich version of unzip do you have (unzip -v, look for ZIP64_SUPPORT)?\nIt seems that (some version of?) OS X ships with an older unzip which\ncan't handle big files:\n\n   https://superuser.com/questions/114011/extract-large-zip-file-50-gb-on-mac-os-x\n\nIs the following check (zip archive with files bigger than 4GB) skipped,\ne.g. because ZIPINFO is missing?  Otherwise I would expect it to fail as\nwell.\n\nRené\n\n"},{"id":"318253","messageId":"e30554f3-1aa3-acea-500b-6392fce902be@web.de","threadId":"45770","inReplyTo":"edf33657-f74b-3cd5-44a7-8e16231bd978@web.de","subject":"Re: [PATCH v3 0/5] archive-zip: support files and archives bigger than 4GB","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2017-04-30T05:31:37Z","receivedAt":"2017-04-30T05:30:59Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"\n\nOn 30/04/17 00:28, René Scharfe wrote:\n> Am 29.04.2017 um 23:00 schrieb Torsten Bögershausen:\n>> This fails here under Mac OS:\n>> commit 4cdf3f9d84568da72f1dcade812de7a42ecb6d15\n>> Author: René Scharfe <l.s.r@web.de>\n>> Date:   Mon Apr 24 19:33:34 2017 +0200\n>>\n>>      archive-zip: support files bigger than 4GB\n>>\n>> ---------------------------\n>> Parts of t5004.log, hope this is helpful:\n>>\n>> \"$GIT_UNZIP\" -t many-big.zip\n>>\n>> Archive:  many-big.zip\n>> warning [many-big.zip]:  577175 extra bytes at beginning or within zipfile\n>>    (attempting to process anyway)\n>> error [many-big.zip]:  start of central directory not found;\n>>    zipfile corrupt.\n>>    (please check that you have transferred or created the zipfile in the\n>>    appropriate BINARY mode and that you have compiled UnZip properly)\n>> not ok 12 - zip archive bigger than 4GB\n>> #\t\n>> #\t\t# build string containing 65536 characters\n>\n> Which version of unzip do you have (unzip -v, look for ZIP64_SUPPORT)?\n> It seems that (some version of?) OS X ships with an older unzip which\n> can't handle big files:\n>\n>    https://superuser.com/questions/114011/extract-large-zip-file-50-gb-on-mac-os-x\n>\n> Is the following check (zip archive with files bigger than 4GB) skipped,\n> e.g. because ZIPINFO is missing?  Otherwise I would expect it to fail as\n> well.\n>\n> René\n>\n\nSorry, I was not looking careful enough, the macro `$GIT_UNZIP`\ngave the impression that an unzip provided by Git (or the Git test framework) \nwas used :-(\n\n$ which unzip\n/usr/bin/unzip\n\n$ unzip -v\nUnZip 5.52 of 28 February 2005, by Info-ZIP.  Maintained by C. Spieler.  Send\nbug reports using http://www.info-zip.org/zip-bug.html; see README for details.\n\nLatest sources and executables are at ftp://ftp.info-zip.org/pub/infozip/ ;\nsee ftp://ftp.info-zip.org/pub/infozip/UnZip.html for other sites.\n\nCompiled with gcc 4.2.1 Compatible Apple LLVM 7.0.0 (clang-700.0.59.1) for Unix \non Aug  1 2015.\n\nUnZip special compilation options:\n         COPYRIGHT_CLEAN (PKZIP 0.9x unreducing method not supported)\n         SET_DIR_ATTRIB\n         TIMESTAMP\n         USE_EF_UT_TIME\n         USE_UNSHRINK (PKZIP/Zip 1.x unshrinking method supported)\n         USE_DEFLATE64 (PKZIP 4.x Deflate64(tm) supported)\n         VMS_TEXT_CONV\n         [decryption, version 2.9 of 05 May 2000]\n\nUnZip and ZipInfo environment options:\n            UNZIP:  [none]\n         UNZIPOPT:  [none]\n          ZIPINFO:  [none]\n       ZIPINFOOPT:  [none]\n\n-------------------\nAnd here is the longer log:\nnot ok 12 - zip archive bigger than 4GB\n#\n#               # build string containing 65536 characters\n# \ns=0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef &&\n# \ns=$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s &&\n# \ns=$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s &&\n#\n#               # create blob with a length of 65536 + 1 bytes\n#               blob=$(echo $s | git hash-object -w --stdin) &&\n#\n#               # create tree containing 65500 entries of that blob\n#               for i in $(test_seq 1 65500)\n#               do\n#                       echo \"100644 blob $blob $i\"\n#               done >tree &&\n#               tree=$(git mktree <tree) &&\n#\n#               # zip it, creating an archive a bit bigger than 4GB\n#               git archive -0 -o many-big.zip $tree &&\n#\n#               \"$GIT_UNZIP\" -t many-big.zip 9999 65500 &&\n#               \"$GIT_UNZIP\" -t many-big.zip\n#\n\nskipping test: zip archive with files bigger than 4GB\n         # Pack created with:\n         #   dd if=/dev/zero of=file bs=1M count=4100 && git hash-object -w file\n         mkdir -p .git/objects/pack &&\n         (\n                 cd .git/objects/pack &&\n                 \"$GIT_UNZIP\" \"$TEST_DIRECTORY\"/t5004/big-pack.zip\n         ) &&\n         blob=754a93d6fada4c6873360e6cb4b209132271ab0e &&\n         size=$(expr 4100 \"*\" 1024 \"*\" 1024) &&\n\n         # create a tree containing the file\n         tree=$(echo \"100644 blob $blob  big-file\" | git mktree) &&\n\n         # zip it, creating an archive with a file bigger than 4GB\n         git archive -o big.zip $tree &&\n\n         \"$GIT_UNZIP\" -t big.zip &&\n         \"$ZIPINFO\" big.zip >big.lst &&\n         grep $size big.lst\n\nok 13 # skip zip archive with files bigger than 4GB (missing ZIPINFO of \nEXPENSIVE,UNZIP,ZIPINFO)\n\n# failed 1 among 13 test(s)\n1..13\n\n\n"},{"id":"318254","messageId":"d8a1edfc-e4d6-2bd2-7b07-a1a10d89490a@web.de","threadId":"45770","inReplyTo":"e30554f3-1aa3-acea-500b-6392fce902be@web.de","subject":"Re: [PATCH v3 0/5] archive-zip: support files and archives bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-30T07:53:52Z","receivedAt":"2017-04-30T07:54:11Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 30.04.2017 um 07:31 schrieb Torsten Bögershausen:\n> Sorry, I was not looking careful enough, the macro `$GIT_UNZIP`\n> gave the impression that an unzip provided by Git (or the Git test \n> framework) was used :-(\n> \n> $ which unzip\n> /usr/bin/unzip\n> \n> $ unzip -v\n> UnZip 5.52 of 28 February 2005, by Info-ZIP.  Maintained by C. Spieler.  \n> Send\n> bug reports using http://www.info-zip.org/zip-bug.html; see README for \n> details.\n> \n> Latest sources and executables are at ftp://ftp.info-zip.org/pub/infozip/ ;\n> see ftp://ftp.info-zip.org/pub/infozip/UnZip.html for other sites.\n> \n> Compiled with gcc 4.2.1 Compatible Apple LLVM 7.0.0 (clang-700.0.59.1) \n> for Unix on Aug  1 2015.\n> \n> UnZip special compilation options:\n>          COPYRIGHT_CLEAN (PKZIP 0.9x unreducing method not supported)\n>          SET_DIR_ATTRIB\n>          TIMESTAMP\n>          USE_EF_UT_TIME\n>          USE_UNSHRINK (PKZIP/Zip 1.x unshrinking method supported)\n>          USE_DEFLATE64 (PKZIP 4.x Deflate64(tm) supported)\n>          VMS_TEXT_CONV\n>          [decryption, version 2.9 of 05 May 2000]\n> \n> UnZip and ZipInfo environment options:\n>             UNZIP:  [none]\n>          UNZIPOPT:  [none]\n>           ZIPINFO:  [none]\n>        ZIPINFOOPT:  [none]\n\nOK, so they indeed still ship the old version of unzip that doesn't\nsupport big files.\n\n> ok 13 # skip zip archive with files bigger than 4GB (missing ZIPINFO of \n> EXPENSIVE,UNZIP,ZIPINFO)\n\nAnd if you had zipinfo then this test would certainly fail.\n\nAnyway, thanks for running these expensive tests!  You could retry\nwith unzip version 6.00 and its zipinfo if you want.  But we\ncertainly need the following patch:\n\n-- >8 --\nSubject: [PATCH] t5004: require 64-bit support for big ZIP tests\n\nCheck if unzip supports the ZIP64 format and skip the tests that create\nbig archives otherwise.  Also skip the test that archives a big file on\n32-bit platforms because the git object systems can't unpack files\nbigger than 4GB there.\n\nReported-by: Torsten Bögershausen <tboegi@web.de>\nSigned-off-by: Rene Scharfe <l.s.r@web.de>\n---\n t/t5004-archive-corner-cases.sh | 9 +++++++--\n 1 file changed, 7 insertions(+), 2 deletions(-)\n\ndiff --git a/t/t5004-archive-corner-cases.sh b/t/t5004-archive-corner-cases.sh\nindex 9106c53c4c..fd23c75f59 100755\n--- a/t/t5004-archive-corner-cases.sh\n+++ b/t/t5004-archive-corner-cases.sh\n@@ -27,6 +27,9 @@ check_dir() {\n \ttest_cmp expect actual\n }\n \n+test_lazy_prereq UNZIP_ZIP64_SUPPORT '\n+\t\"$GIT_UNZIP\" -v | grep ZIP64_SUPPORT\n+'\n \n # bsdtar/libarchive versions before 3.1.3 consider a tar file with a\n # global pax header that is not followed by a file record as corrupt.\n@@ -155,7 +158,8 @@ test_expect_success ZIPINFO 'zip archive with many entries' '\n \ttest_cmp expect actual\n '\n \n-test_expect_success EXPENSIVE,UNZIP 'zip archive bigger than 4GB' '\n+test_expect_success EXPENSIVE,UNZIP,UNZIP_ZIP64_SUPPORT \\\n+\t'zip archive bigger than 4GB' '\n \t# build string containing 65536 characters\n \ts=0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef &&\n \ts=$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s$s &&\n@@ -178,7 +182,8 @@ test_expect_success EXPENSIVE,UNZIP 'zip archive bigger than 4GB' '\n \t\"$GIT_UNZIP\" -t many-big.zip\n '\n \n-test_expect_success EXPENSIVE,UNZIP,ZIPINFO 'zip archive with files bigger than 4GB' '\n+test_expect_success EXPENSIVE,LONG_IS_64BIT,UNZIP,UNZIP_ZIP64_SUPPORT,ZIPINFO \\\n+\t'zip archive with files bigger than 4GB' '\n \t# Pack created with:\n \t#   dd if=/dev/zero of=file bs=1M count=4100 && git hash-object -w file\n \tmkdir -p .git/objects/pack &&\n-- \n2.12.2\n"},{"id":"318257","messageId":"1448ad65-126a-41a0-cc1c-49d54aed3f26@web.de","threadId":"45770","inReplyTo":"d8a1edfc-e4d6-2bd2-7b07-a1a10d89490a@web.de","subject":"Re: [PATCH v3 0/5] archive-zip: support files and archives bigger than 4GB","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2017-04-30T13:06:45Z","receivedAt":"2017-04-30T13:07:00Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 2017-04-30 09:53, René Scharfe wrote:\n> Am 30.04.2017 um 07:31 schrieb Torsten Bögershausen:\n>> Sorry, I was not looking careful enough, the macro `$GIT_UNZIP`\n>> gave the impression that an unzip provided by Git (or the Git test \n>> framework) was used :-(\n>>\n>> $ which unzip\n>> /usr/bin/unzip\n>>\n>> $ unzip -v\n>> UnZip 5.52 of 28 February 2005, by Info-ZIP.  Maintained by C. Spieler.  \n>> Send\n>> bug reports using http://www.info-zip.org/zip-bug.html; see README for \n>> details.\n>>\n>> Latest sources and executables are at ftp://ftp.info-zip.org/pub/infozip/ ;\n>> see ftp://ftp.info-zip.org/pub/infozip/UnZip.html for other sites.\n>>\n>> Compiled with gcc 4.2.1 Compatible Apple LLVM 7.0.0 (clang-700.0.59.1) \n>> for Unix on Aug  1 2015.\n>>\n>> UnZip special compilation options:\n>>          COPYRIGHT_CLEAN (PKZIP 0.9x unreducing method not supported)\n>>          SET_DIR_ATTRIB\n>>          TIMESTAMP\n>>          USE_EF_UT_TIME\n>>          USE_UNSHRINK (PKZIP/Zip 1.x unshrinking method supported)\n>>          USE_DEFLATE64 (PKZIP 4.x Deflate64(tm) supported)\n>>          VMS_TEXT_CONV\n>>          [decryption, version 2.9 of 05 May 2000]\n>>\n>> UnZip and ZipInfo environment options:\n>>             UNZIP:  [none]\n>>          UNZIPOPT:  [none]\n>>           ZIPINFO:  [none]\n>>        ZIPINFOOPT:  [none]\n> \n> OK, so they indeed still ship the old version of unzip that doesn't\n> support big files.\n> \n>> ok 13 # skip zip archive with files bigger than 4GB (missing ZIPINFO of \n>> EXPENSIVE,UNZIP,ZIPINFO)\n> \n> And if you had zipinfo then this test would certainly fail.\n> \n\nAfter installing unzip via mac ports, I have:\n\n$ which zipinfo\n/opt/local/bin/zipinfo\n\n$ which unzip\n /opt/local/bin/unzip\n\n$ unzip -v\nUnZip 6.00 of 20 April 2009, by Info-ZIP.  Maintained by C. Spieler.  Send\nbug reports using http://www.info-zip.org/zip-bug.html; see README for details.\n\nLatest sources and executables are at ftp://ftp.info-zip.org/pub/infozip/ ;\nsee ftp://ftp.info-zip.org/pub/infozip/UnZip.html for other sites.\n\nCompiled with gcc 4.2.1 Compatible Apple LLVM 7.0.2 (clang-700.1.81) for Unix\nMac OS X on Jan 31 2016.\n\nUnZip special compilation options:\n        COPYRIGHT_CLEAN (PKZIP 0.9x unreducing method not supported)\n        SET_DIR_ATTRIB\n        SYMLINKS (symbolic links supported, if RTL and file system permit)\n        TIMESTAMP\n        UNIXBACKUP\n        USE_EF_UT_TIME\n        USE_UNSHRINK (PKZIP/Zip 1.x unshrinking method supported)\n        USE_DEFLATE64 (PKZIP 4.x Deflate64(tm) supported)\n        VMS_TEXT_CONV\n        [decryption, version 2.11 of 05 Jan 2007]\n\nUnZip and ZipInfo environment options:\n           UNZIP:  [none]\n        UNZIPOPT:  [none]\n         ZIPINFO:  [none]\n      ZIPINFOOPT:  [none]\n\n\n> Anyway, thanks for running these expensive tests!  You could retry\n> with unzip version 6.00 and its zipinfo if you want.  But we\n> certainly need the following patch:\n> \n\nThe 6.0 is not compiled with ZIP64_SUPPORT....\nAfter manually doing a copy-paste of your patch from below, tests are skipped:\n\nok 11 - zip archive with many entries\nok 12 # skip zip archive bigger than 4GB (missing UNZIP_ZIP64_SUPPORT of\nEXPENSIVE,UNZIP,UNZIP_ZIP64_SUPPORT)\nok 13 # skip zip archive with files bigger than 4GB (missing UNZIP_ZIP64_SUPPORT\nof EXPENSIVE,UNZIP,UNZIP_ZIP64_SUPPORT,ZIPINFO)\n\n[patch snipped]\n\n"},{"id":"318261","messageId":"9f6cb421-db61-51ca-6a4b-ea7c94bd513e@kdbg.org","threadId":"45770","inReplyTo":"d8a1edfc-e4d6-2bd2-7b07-a1a10d89490a@web.de","subject":"Re: [PATCH v3 0/5] archive-zip: support files and archives bigger than 4GB","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2017-04-30T16:32:23Z","receivedAt":"2017-04-30T16:32:32Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 30.04.2017 um 09:53 schrieb René Scharfe:\n> @@ -178,7 +182,8 @@ test_expect_success EXPENSIVE,UNZIP 'zip archive bigger than 4GB' '\n>  \t\"$GIT_UNZIP\" -t many-big.zip\n>  '\n>\n> -test_expect_success EXPENSIVE,UNZIP,ZIPINFO 'zip archive with files bigger than 4GB' '\n> +test_expect_success EXPENSIVE,LONG_IS_64BIT,UNZIP,UNZIP_ZIP64_SUPPORT,ZIPINFO \\\n\nWhy is LONG_IS_64BIT required?\n\n> +\t'zip archive with files bigger than 4GB' '\n>  \t# Pack created with:\n>  \t#   dd if=/dev/zero of=file bs=1M count=4100 && git hash-object -w file\n>  \tmkdir -p .git/objects/pack &&\n>\n\n-- Hannes\n\n"},{"id":"318262","messageId":"d1cea1c3-f974-c839-dbdd-5bc95756be84@web.de","threadId":"45770","inReplyTo":"9f6cb421-db61-51ca-6a4b-ea7c94bd513e@kdbg.org","subject":"Re: [PATCH v3 0/5] archive-zip: support files and archives bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-04-30T16:40:00Z","receivedAt":"2017-04-30T16:40:14Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 30.04.2017 um 18:32 schrieb Johannes Sixt:\n> Am 30.04.2017 um 09:53 schrieb René Scharfe:\n>> @@ -178,7 +182,8 @@ test_expect_success EXPENSIVE,UNZIP 'zip archive \n>> bigger than 4GB' '\n>>      \"$GIT_UNZIP\" -t many-big.zip\n>>  '\n>>\n>> -test_expect_success EXPENSIVE,UNZIP,ZIPINFO 'zip archive with files \n>> bigger than 4GB' '\n>> +test_expect_success \n>> EXPENSIVE,LONG_IS_64BIT,UNZIP,UNZIP_ZIP64_SUPPORT,ZIPINFO \\\n> \n> Why is LONG_IS_64BIT required?\n\nBlob sizes are kept in variables of type unsigned long.  64 bits are\nrequired to store file sizes bigger than 4GB, and this test is about\nsuch a file.  A 32-bit git can't use the pack we supply the test file\nin, so we have to skip this test.\n\n>> +    'zip archive with files bigger than 4GB' '\n>>      # Pack created with:\n>>      #   dd if=/dev/zero of=file bs=1M count=4100 && git hash-object \n>> -w file\n>>      mkdir -p .git/objects/pack &&\n>>\n> \n> -- Hannes\n> \n"},{"id":"318272","messageId":"xmqqfugplamn.fsf@gitster.mtv.corp.google.com","threadId":"45770","inReplyTo":"d1cea1c3-f974-c839-dbdd-5bc95756be84@web.de","subject":"Re: [PATCH v3 0/5] archive-zip: support files and archives bigger than 4GB","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-04-30T23:49:04Z","receivedAt":"2017-04-30T23:49:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"René Scharfe <l.s.r@web.de> writes:\n\n> Am 30.04.2017 um 18:32 schrieb Johannes Sixt:\n>> Am 30.04.2017 um 09:53 schrieb René Scharfe:\n>>> @@ -178,7 +182,8 @@ test_expect_success EXPENSIVE,UNZIP 'zip\n>>> archive bigger than 4GB' '\n>>>      \"$GIT_UNZIP\" -t many-big.zip\n>>>  '\n>>>\n>>> -test_expect_success EXPENSIVE,UNZIP,ZIPINFO 'zip archive with\n>>> files bigger than 4GB' '\n>>> +test_expect_success\n>>> EXPENSIVE,LONG_IS_64BIT,UNZIP,UNZIP_ZIP64_SUPPORT,ZIPINFO \\\n>>\n>> Why is LONG_IS_64BIT required?\n>\n> Blob sizes are kept in variables of type unsigned long.  64 bits are\n> required to store file sizes bigger than 4GB, and this test is about\n> such a file.  A 32-bit git can't use the pack we supply the test file\n> in, so we have to skip this test.\n>\n>>> +    'zip archive with files bigger than 4GB' '\n>>>      # Pack created with:\n>>>      #   dd if=/dev/zero of=file bs=1M count=4100 && git\n>>> hash-object -w file\n>>>      mkdir -p .git/objects/pack &&\n\nOK, so let's queue this on top and have it in 'next' to unblock\nusers of older unzip and unzip compiled wihtout zip64 support?\n\nThanks.\n\n\n"},{"id":"318354","messageId":"85744f48-c408-4fc7-bdf1-1e3961703dd0@web.de","threadId":"45770","inReplyTo":"xmqqfugplamn.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v3 0/5] archive-zip: support files and archives bigger than 4GB","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-05-01T08:30:04Z","receivedAt":"2017-05-01T08:30:33Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 01.05.2017 um 01:49 schrieb Junio C Hamano:\n> René Scharfe <l.s.r@web.de> writes:\n> \n>> Am 30.04.2017 um 18:32 schrieb Johannes Sixt:\n>>> Am 30.04.2017 um 09:53 schrieb René Scharfe:\n>>>> @@ -178,7 +182,8 @@ test_expect_success EXPENSIVE,UNZIP 'zip\n>>>> archive bigger than 4GB' '\n>>>>       \"$GIT_UNZIP\" -t many-big.zip\n>>>>   '\n>>>>\n>>>> -test_expect_success EXPENSIVE,UNZIP,ZIPINFO 'zip archive with\n>>>> files bigger than 4GB' '\n>>>> +test_expect_success\n>>>> EXPENSIVE,LONG_IS_64BIT,UNZIP,UNZIP_ZIP64_SUPPORT,ZIPINFO \\\n>>>\n>>> Why is LONG_IS_64BIT required?\n>>\n>> Blob sizes are kept in variables of type unsigned long.  64 bits are\n>> required to store file sizes bigger than 4GB, and this test is about\n>> such a file.  A 32-bit git can't use the pack we supply the test file\n>> in, so we have to skip this test.\n>>\n>>>> +    'zip archive with files bigger than 4GB' '\n>>>>       # Pack created with:\n>>>>       #   dd if=/dev/zero of=file bs=1M count=4100 && git\n>>>> hash-object -w file\n>>>>       mkdir -p .git/objects/pack &&\n> \n> OK, so let's queue this on top and have it in 'next' to unblock\n> users of older unzip and unzip compiled wihtout zip64 support?\n\nYes, please; it allows them to run t5004 with --long-tests or \nGIT_TEST_LONG set (but not to actually test zip64 functionality).\n\nRené\n"}]}