libxml2

c/libxml2

mirror of https://gitlab.gnome.org/GNOME/libxml2 synced 2025-03-28 21:33:13 +00:00

Author	SHA1	Message	Date
Nick Wellnhofer	fd1b939168	include: Convert some macros to enums	2025-03-14 00:35:40 +01:00
Nick Wellnhofer	69b83bb68e	encoding: Detect truncated multi-byte sequences with ICU Unlike iconv or the internal converters, ICU consumes truncated multi- byte sequences at the end of an input buffer. We currently check for a non-empty raw input buffer to detect truncated sequences, so this fails with ICU. It might be possible to inspect the pivot buffer pointers, but it seems cleaner to implement a `flush` flag for some encoding and I/O functions. After flushing, we can check for U_TRUNCATED_CHAR_FOUND with ICU, or detect remaining input with other converters. Also fix detection of truncated sequences for HTML, XML content and DTDs with iconv.	2025-03-13 22:15:10 +01:00
Nick Wellnhofer	361f7bff92	parser: Make nodePush, nodePop, namePush, namePop private	2025-03-04 16:47:14 +01:00
Nick Wellnhofer	9c16a153d8	Revert "include: Make most IS_* macros private" This reverts commit 84a6c82ff83d04963d6e1c5cd18ded68ea02d99f.	2025-02-13 20:20:17 +01:00
Nick Wellnhofer	c134e8b4dc	include: Make INPUT_CHUNK macro private	2024-12-21 20:02:34 +01:00
Nick Wellnhofer	84a6c82ff8	include: Make most IS_* macros private Macros like IS_DIGIT or IS_LETTER severely pollute the C namespace.	2024-12-21 20:01:30 +01:00
Nick Wellnhofer	0dd910e82b	save: Fix handling of catastrophic errors Don't overwrite catastrophic errors xmlSaveErr. Overwrite non-catastrophic errors in xmlOutputBufferClose.	2024-12-19 02:30:36 +01:00
Nick Wellnhofer	57087e5fc7	parser: Don't overwrite catastrophic errors Stop reporting errors after a catastrophic error. Also make sure that ctxt->errNo matches ctxt->lastError.code.	2024-11-26 00:47:48 +01:00
Nick Wellnhofer	0bc4608c50	html: Use hash table to check for duplicate attributes	2024-10-06 20:04:00 +02:00
Nick Wellnhofer	8af55c8d20	parser: Rename new input API functions These weren't made public yet.	2024-07-11 01:33:29 +02:00
Nick Wellnhofer	d74ca59491	parser: Rename internal xmlNewInput functions	2024-07-11 01:31:50 +02:00
Nick Wellnhofer	38195cf596	parser: Don't produce names with invalid UTF-8 in recovery mode	2024-07-06 15:33:06 +02:00
Nick Wellnhofer	16e7ecd478	xinclude: Check URI length Don't report long URIs as OOM errors.	2024-07-01 18:03:06 +02:00
Nick Wellnhofer	5238404325	parser: Pass resource type to resource loader	2024-06-12 16:36:12 +02:00
Nick Wellnhofer	89fcae4dfd	parser: Don't report malloc failures when creating context We don't want messages to stderr before an error handler could be set on a parser context.	2024-06-12 16:36:12 +02:00
Nick Wellnhofer	b9d2f3c911	parser: Introduce new input API - xmlInputCreateUrl - xmlInputCreateMemory - xmlInputCreateString - xmlInputCreateFd - xmlInputCreateIO - xmlInputSetEncoding These functions don't take a parser context and work on xmlParserInputs, replacing functions working on xmlParserInputBuffers. xmlInputCreateUrl and xmlInputSetEncoding offer fine-grained error handling. Several XML_INPUT_* flags offer additional control.	2024-06-12 16:22:52 +02:00
Nick Wellnhofer	ff3b091910	parser: Implement XML_PARSE_NO_UNZIP option	2024-06-12 16:14:15 +02:00
Nick Wellnhofer	b47a95fe31	parser: Don't make xmlCtxtErrIO public	2024-05-20 14:22:56 +02:00
Nick Wellnhofer	84a71860a8	xmlreader: Fix xmlTextReaderConstEncoding Regression from commit f1c1f5c6. Fixes #697.	2024-02-26 15:33:06 +01:00
Nick Wellnhofer	8961056f9b	parser: Make experimental input API private This needs to be reworked.	2024-01-23 00:47:44 +01:00
Nick Wellnhofer	37c6618be5	parser: Rework parsing of attribute and entity values Don't use a separate function to handle "complex" attributes. Validate UTF-8 byte sequences without decoding. This should improve performance considerably when parsing multi-byte UTF-8 sequences. Use a string buffer to avoid unnecessary allocations and copying when expanding entities. Normalize attribute values in a single pass while expanding entities. Be more lenient in recovery mode. If no entity substitution was requested, validate entities without expanding. Fixes #596. Also fixes #655.	2024-01-02 15:42:03 +01:00
Nick Wellnhofer	7e0bbbc143	parser: New input API Provide a new set of functions to create xmlParserInputs. These can be used for the document entity or from external entity loaders. - Don't require xmlParserInputBuffer. - All functions take a base URI. - All functions take an encoding as string. - xmlNewInputURL also takes a public ID. - xmlNewInputMemory takes a size_t. - Optimization hints for memory buffers. Improve documentation. Only call xmlInitParser before allocating a new parser context. Call xmlCtxtUseOptions as early as possible.	2023-12-29 01:22:13 +01:00
Nick Wellnhofer	6a9a88a17f	parser: Move progressive flag into input struct	2023-12-29 01:20:08 +01:00
Nick Wellnhofer	d944a41515	parser: Fix in-parameter-entity and in-external-dtd checks Use in ctxt->input->entity instead of ctxt->inputNr to determine whether we are inside a parameter entity. Stop using ctxt->external to check whether we're in an external DTD. This is signaled by ctxt->inSubset == 2.	2023-12-29 01:19:56 +01:00
Nick Wellnhofer	130436917c	parser: Rename xmlErrParser to xmlCtxtErr	2023-12-21 15:02:24 +01:00
Nick Wellnhofer	8d0aaf4b95	parser: Remove xmlErrEncoding Use xmlFatalErr or xmlCtxtErrIO.	2023-12-21 15:02:24 +01:00
Nick Wellnhofer	54c70ed57f	parser: Improve error handling Introduce xmlCtxtSetErrorHandler allowing to set a structured error for a parser context. There already was the "serror" SAX handler but this always receives the parser context as argument. Start to use xmlRaiseMemoryError. Remove useless arguments from memory error functions. Rename xmlErrMemory to xmlCtxtErrMemory. Remove a few calls to xmlGenericError. Remove support for runtime entity debugging.	2023-12-21 02:46:27 +01:00
Nick Wellnhofer	f19a95108a	parser: Report malloc failures Fix many places where malloc failures aren't reported. Make xmlErrMemory public. This is useful for custom external entity loaders. Introduce new API function xmlSwitchEncodingName. Change the way how we store whether the the parser is stopped. This used to be signaled by setting ctxt->instate to XML_PARSER_EOF which was misdesigned and error-prone. Set ctxt->disableSAX to 2 instead and introduce a macro PARSER_STOPPED. Also stop to remove parser inputs in xmlHaltParser. This allows to remove many checks of ctxt->instate. Introduce xmlErrParser to handle errors if a parser context is available.	2023-12-11 22:13:05 +01:00
Nick Wellnhofer	c082ef4644	parser: Stop switching to ISO-8859-1 on encoding errors Use U+FFFD Replacement Character if invalid UTF-8 is encountered in recovery mode. Also rewrite xmlNextChar and xmlCurrentChar. Fixes #598.	2023-10-22 16:32:54 +02:00
Nick Wellnhofer	eb69c1d39d	parser: Fix initialization of namespace data Move initialization to xmlInitSAXParserCtxt. Also add missing XML_HIDDEN to xmlParserNsFree. Fixes #597.	2023-10-02 12:33:29 +02:00
Nick Wellnhofer	e0dd330b8f	parser: Use hash tables to avoid quadratic behavior Use a hash table to lookup namespaces by prefix. The hash table stores an index into the namespace table. Auxiliary data for namespaces is stored in a separate array along the main namespace table. Use a hash table to verify attribute uniqueness. The hash table stores an index into the attribute table. Reuse hash value from the dictionary to avoid computing them twice. See #346.	2023-09-29 12:43:22 +02:00
Nick Wellnhofer	f1c1f5c6b4	parser: Revert change to doc->encoding Fixes #579.	2023-08-17 12:47:14 +02:00
Nick Wellnhofer	ec7be50662	parser: Rework encoding detection Introduce XML_INPUT_HAS_ENCODING flag for xmlParserInput which is set when xmlSwitchEncoding is called. The parser can use the flag to reliably detect whether an encoding was already set via user override, BOM or other auto-detection. In this case, the encoding declaration won't be used to switch the encoding. Before, an inscrutable mix of ctxt->charset, ctxt->input->encoding and ctxt->input->buf->encoder was used. Introduce private helper functions to switch encodings used by both the XML and HTML parser: - xmlDetectEncoding which skips over the BOM, allowing to remove the BOM checks from other encoding functions. - xmlSetDeclaredEncoding, replacing htmlCheckEncodingDirect, which warns about encoding mismatches. If users override the encoding, store the declared instead of the actual encoding in xmlDoc. In this case, the actual encoding is known and the raw value from the doc is more useful. Also use the input flags to store the ISO-8859-1 fallback state. Restrict the fallback to cases where no encoding was specified. (The fallback is only useful in recovery mode and these days broken UTF-8 is probably more likely than ISO-8859-1, so it might eventually be removed completely.) The 'charset' member of xmlParserCtxt is now unused. The 'encoding' member of xmlParserInput is now unused. The 'standalone' member of xmlParserInput is renamed to 'flags'. A new parser state XML_PARSER_XML_DECL is added for the push parser.	2023-08-08 15:19:46 +02:00
Nick Wellnhofer	fc69cf568b	parser: Move xmlFatalErr to parserInternals.c	2023-04-30 17:51:29 +02:00
Nick Wellnhofer	04d1bedd8c	parser: Rework shrinking of input buffers Don't try to grow the input buffer in xmlParserShrink. This makes sure that no memory allocations are made and the function always succeeds. Remove unnecessary invocations of SHRINK. Invoke SHRINK at the end of DTD parsing loops. Shrink before growing.	2023-03-21 13:19:18 +01:00
Nick Wellnhofer	b167c73144	parser: Fix short-lived regression causing infinite loops Fix 3eb6bf03. We really have to halt the parser, so the input buffer gets reset.	2023-03-14 15:16:04 +01:00
Nick Wellnhofer	2099441f32	parser: Stop calling xmlParserInputShrink Introduce xmlParserShrink which takes a parser context to simplify error handling.	2023-03-13 17:51:13 +01:00
Nick Wellnhofer	3eb6bf0386	parser: Stop calling xmlParserInputGrow Introduce xmlParserGrow which takes a parser context to simplify error handling.	2023-03-12 17:05:51 +01:00
Nick Wellnhofer	ccb6d54409	Hide internal functions These functions were never declared in public headers, so it should be safe to hide them. Fixes #139.	2022-11-27 02:20:53 +01:00
Nick Wellnhofer	0f568c0b73	Consolidate private header files Private functions were previously declared - in header files in the root directory - in public headers guarded with IN_LIBXML - in libxml.h - redundantly in source files that used them. Consolidate all private header files in include/private.	2022-08-26 02:11:56 +02:00

40 Commits